Dialogue is the fastest way to see what a language model does to writing, because dialogue is made of the things a most-probable continuation cannot supply: withheld information, mismatched intentions, and people who talk past each other. Here are five failures, each shown rather than asserted.
Why dialogue exposes the machine fastest
Most prose can be competent and generic at once. Dialogue cannot, because the reader is running a model of each speaker and notices immediately when two of them are the same person. Every weakness that is subtle in exposition is loud in a scene.
The examples below are written in the register generated dialogue tends to land in. They are not transcripts of any particular tool’s output; they are constructed to isolate one failure each, which is what makes the repair legible. The same five properties are what the craft argument about fiction rests on, in a form you can see rather than argue about.
One: one voice, several mouths
Every character uses the same sentence length, the same register, the same vocabulary, and the same willingness to explain themselves.
FLAT
"I think we should reconsider the timeline," said Maya. "The engineering
team has raised legitimate concerns about the deployment schedule."
"I understand your concerns," Tom replied, "but we've already committed
to the client. Moving the date would damage the relationship."
"I appreciate that, but shipping something broken would damage it more."
REPAIRED
"Push it," Maya said.
"To when?"
"I don't know yet. Ask Priya, she's the one who has to be awake at four."
Tom looked at the calendar and did not write anything on it. "I told
Dettmer the fourteenth."
"You told Dettmer the fourteenth."
"That's what I said."
"No, I'm agreeing with you," Maya said. "You told him. Not us."
Enter fullscreen mode Exit fullscreen mode
What changed: unequal sentence lengths, one character who answers questions and one who does not, a repeated line that means something different the second time, and a piece of information (Priya, four in the morning) that exists to characterise rather than to inform. The flat version has none of that because none of it is the likeliest continuation.
Two: everybody says what they mean
Generated characters are unfailingly articulate about their own emotional states. Real dialogue is mostly people talking about something adjacent to the thing.
FLAT
"I feel like you don't value my contributions to this project," she said.
"That's not true. I value you enormously. I'm just under a lot of pressure
right now and I haven't been showing it well."
REPAIRED
"You changed the ordering," she said.
"It reads better."
"It reads like you wrote it."
He started to say something and then went back to the screen. "I can put
it back."
"Don't put it back." She picked up her coat. "Just tell me next time, so
I know before Anders does."
Enter fullscreen mode Exit fullscreen mode
Nobody names an emotion and both are visible. The subject on the surface is the ordering of a document; the subject underneath is standing. That gap is subtext, and it is a deliberate withholding — which is exactly the operation a model optimising for the most probable next token will not perform, because the probable thing to say is the thing you mean.
Three: information delivered, not withheld
Characters tell each other things they both already know, because the reader needs them and generated dialogue serves the reader directly rather than through the characters.
FLAT
"As you know, our father left the mill to both of us when he died three
years ago, and since then we've been struggling to keep it profitable."
"Yes, and now the bank is threatening to foreclose unless we can make the
payment by the end of the month."
REPAIRED
"How much."
"You know how much."
"I want to hear you say it."
He said the number. It came out flatter than he meant it to, and she
laughed, once, with no pleasure in it at all.
"Dad would have paid it out of the float," she said.
"Dad had a float."
Enter fullscreen mode Exit fullscreen mode
The repair gives the reader less and the scene more. Both the amount and the family history stay off the page; what arrives instead is that she is angry, he is ashamed, and the father was better at this. A model asked to write this scene will supply the number, because the instruction implies the reader needs it and supplying is the higher-probability move.
Four: the uniform beat
Line, attribution, small action, line, attribution, small action. The rhythm is regular to the point of metronomic, and it flattens emphasis because a beat that appears everywhere emphasises nothing.
FLAT
"We should go," he said, glancing at the door.
"Not yet," she replied, setting down her cup.
"They're waiting," he insisted, shifting his weight.
"Let them wait," she answered, meeting his eyes.
REPAIRED
"We should go."
"Not yet."
"They're waiting."
She put the cup down. It took her a while to do it.
"Let them wait."
Enter fullscreen mode Exit fullscreen mode
Four beats removed and one kept. The one that survives now carries everything, because it is the only physical action in the passage. The rule generalises: beats are punctuation, and prose that punctuates every line has no punctuation at all.
Five: the scene resolves
Generated scenes reach agreement. Someone concedes, someone acknowledges the other’s point, and the conversation lands somewhere stable — which is the death of a plot, because a scene that resolves has no reason to be followed by another one.
FLAT
"Maybe you're right," he said finally. "I've been too focused on the
numbers. Let's do it your way and see what happens."
She smiled. "Thank you. That means a lot."
REPAIRED
"Fine," he said. "Your way."
"That isn't what I asked for."
"It's what you're getting."
She thought about it for long enough that he started to gather his
papers, and then she said: "I'll need it in writing," and he stopped.
Enter fullscreen mode Exit fullscreen mode
The repaired scene ends with a new problem rather than a settled one. This failure is the hardest of the five to prompt away, because resolution is what conversations in text overwhelmingly do — transcripts, forum threads, screenplays with the scene breaks removed. The prior is enormous.
One mechanism behind all five
Every failure above is the same property in a different costume. The model produces the most probable continuation given the context; good dialogue is a sequence of improbable continuations that turn out, once you have read them, to have been the right ones.
- Distinct voices are improbable. Given a scene about a deadline, the likeliest next line is the one a competent professional would say, and every character in the scene qualifies as one.
- Subtext is a withheld token. Saying the adjacent thing is by construction less likely than saying the thing.
- Withholding exposition fights the instruction. The prompt said write a scene about the foreclosure, so the foreclosure is what the tokens are about.
- Uniform rhythm is the mean. Averaged over all written dialogue, an attribution and a small gesture is the modal beat.
- Resolution is the modal ending.
Turning the temperature up does not fix this. It makes the improbable more likely in general, which produces stranger word choices rather than better withholding — the choices that matter here are structural, not lexical, and sampling parameters operate on single tokens. What does help slightly is giving each character a written constraint they cannot violate (what they will not admit, what they always deflect to, how many words they use), because a constraint reshapes the distribution instead of flattening it.
Where it does help
- Diagnosis of your own dialogue. “Strip the attributions from this scene and tell me who is speaking each line” is a real test with a real answer, and if it cannot tell, neither can your reader.
- Consistency checking. A character who used “shan’t” in chapter two and “won’t gonna” in chapter nine is a mechanical error over long text, and machines are better than people at long text.
- Register research. How a particular trade actually talks about its equipment, what an era’s idiom sounded like, what an occupation calls a tool. Verify it like any other fact, because it is one.
- The bad draft to react against. Genuinely useful: generate the flat version deliberately, then write the line the character would say instead. Reacting is easier than starting, and the flat version costs nothing to throw away.
- Cleaning up dictation. If you speak your dialogue, punctuating and paragraphing a transcript is a mechanical job that does not touch what was said.
- Matching your own narrative register in the prose around the dialogue, where exemplars from your own archive transfer surface reliably — and where the limits of surface transfer matter far less than they do in the speech itself.
답글 남기기