← LEFT BRAIN DOMINANT

The Mechanics of the Setup-Punchline Delay

Jan 21, 2024 · 5 MIN READ
See also, the right brain on this: Why the Third One Is Always Funnier

Comic timing is an expectation-violation curve with a narrow optimum. That sentence is not funny, which is the whole problem with the subject: you can dissect it, but as E.B. White said, the frog tends to die in the process. I am going to dissect it anyway, because I think the mechanism is actually clarifying, and because once you see it you cannot stop seeing it.

The brain is a predictor

Start with the model that has quietly taken over cognitive science: the brain is a prediction machine. At every level, from retinal cells to language comprehension, it's constantly generating a forecast of what comes next and comparing it to what arrives. The difference, the prediction error, is the signal that gets processed. Confirmed predictions are cheap and mostly ignored. Errors are expensive and get attention.

This is why you can read a sentence with a missing word and not notice, and why a wrong note in a familiar song is jarring. You were predicting. The prediction is most of the perception.

A joke is a device for manufacturing a specific kind of prediction error on purpose.

Setup, then violation

The setup does two jobs. It builds a model in the listener's head, a small world with rules, and it builds it fast enough that the listener commits to a prediction without noticing they have made one. The punchline then delivers something that's inconsistent with that model but consistent with a different, previously unconsidered one.

The formal version of this is called incongruity-resolution, and it has been the leading theory of verbal humor for about fifty years. The incongruity is the prediction error. The resolution is the moment you find the second model that makes the punchline make sense. Both are required. Pure incongruity is just confusion. Pure resolution is just information. A joke is the two in quick succession: the floor drops, and then, immediately, there's a different floor.

Peter McGraw's benign-violation theory adds the third constraint: the violation has to be safe. The model that got broken must be one you weren't depending on. Which is why a joke about your job lands at a party and doesn't land in a performance review, and why timing on dark material is so much harder. The safety margin is thinner.

Where the delay goes

Now the pause. The delay between setup and punchline isn't decoration. It's doing computational work, and the work has a specific optimum.

The listener needs time to build the model. Not consciously. Predictive processing runs on the order of hundreds of milliseconds for a sentence, and the setup's model has to be complete and committed before the punchline arrives, or the violation has nothing to violate. Deliver the punchline too early and you are contradicting a prediction the listener has not finished making. The error signal is weak. The joke reads as a non sequitur.

But the listener is also a predictor, and if you give them too much time, they start predicting the punchline. A pause that runs long shifts the listener from "building the setup model" to "wondering what comes next," and once they're wondering, they're generating candidates, and if one of those candidates is your punchline, the error signal on delivery is zero. They saw it coming. The joke dies of anticipation.

So the window is: after the model is built, before the search for the punchline begins. Empirically, for a spoken one-liner, that's somewhere between about half a second and a second and a half, and it is different for every joke and every room, which is why timing is a skill and not a setting.

The rule of three

The rule of three is this mechanism at its most transparent, and it explains why the third item is the funny one and not the second or the fourth.

One item is a data point. Two items are a pattern. That's the minimum number of points needed to establish a trend, and once two items share a structure, the predictive brain fits a line through them and extrapolates. The third item is the first one that the listener has a confident prediction for. Break the pattern there and the prediction error is maximal.

Why not the second? Because there's no pattern yet to break. A single item doesn't establish a rule, so a second item that deviates isn't a violation, it's just a second item.

Why not the fourth? Because by the fourth, the listener has been given three, has fit the pattern with high confidence, and, crucially, has had time to start predicting the break. The rule of three is famous. Listeners know it. A fourth item is arriving into a room that has already braced. The surprise is spent.

Three is the smallest number that produces a confident prediction, and the largest number before the listener starts predicting the violation itself. It's a narrow optimum in the count, the same way the pause is a narrow optimum in time.

Surprise and inevitability

The best description I know of a good punchline is that it should be surprising and, in retrospect, inevitable. In information-theoretic terms: low probability given the setup as the listener modeled it, high probability given the setup as it actually was.

That second half is what separates a joke from a random swerve. The punchline has to be latent in the setup. The information was there, the listener just built the wrong model from it. When the punchline lands, the listener doesn't simply see the new model, they see that the old model was their own mistake, that the setup was honest the whole time. The laugh is partly at the joke and partly at yourself for having been so confidently wrong about something that was, it turns out, right in front of you.

This is why joke structure is so brittle. Move one word in the setup and the listener builds a different model, and the punchline is either obvious or unrelated. The setup, in other words, was never a preamble. It's a carefully worded misdirection that has to stay technically true.

What laughter is for

The last piece is the noise. Why a physical, involuntary, contagious, audible response to a prediction error?

The leading guess is that laughter predates language and started as a social signal, a way of broadcasting to the group that an apparent threat wasn't real. The play-panting of apes. What it says is roughly: I predicted wrong, and I'm fine, and you can be fine too. That's the benign in benign violation, made audible. It's why laughter is so much stronger in groups, why it spreads, and why a joke told to an empty room isn't really a joke, it's a proposal.

Which, to close the loop, is why timing is the whole art. A joke is a bet that you can build a model in a stranger's head, get them to commit to it without knowing they have, break it at the precise moment their commitment peaks, hand them the replacement before they can feel lost, and do all of this in a way that's safe enough that their nervous system broadcasts relief instead of alarm. The delay is where the bet is placed.

Anyway, a man walks into a bar. That's the setup. You have already built the model. Notice that you're waiting.