Here is the failure mode that has haunted synthetic narration since it existed: every line is rendered alone. A character asks a sharp question, the reply comes back in a neutral register, and the exchange lands as two separate recordings played back to back rather than as a conversation. You can hear the seam even when you can't name it.
The standard workaround is instruction. Tell the engine how each line should sound — impatient, softly, with growing alarm — and stitch the performance together from directions. It works, in the sense that it produces variation. But it also means the performance is coming from a label rather than from the story, and labels are blunt. As of this week, Midsummerr productions no longer work that way.
Every line now hears the lines before it
Each segment we synthesize carries up to five preceding segments with it as context. The engine renders your line already knowing how the scene arrived at it — who just spoke, at what length, in what shape.
The effect is most obvious in fast exchanges. Quick turn-taking used to drag: short lines came back padded and oddly weighted, because nothing told the renderer that this was a clipped back-and-forth rather than a standalone sentence. With the preceding lines attached, the pacing of a rapid dialogue scene comes out of the dialogue. Emotional carry-over works the same way — a reply following a raised voice inherits the temperature of the moment instead of resetting to neutral.
This applies on every path, which matters more than it sounds. A retry after a failed render, a retake triggered by our quality check, a single line you regenerate by hand — all of them send the same context as the original pass. A re-rendered line therefore matches the take sitting next to it, instead of arriving as the one segment that sounds subtly detached from its own scene.
Ready to try it on your own book?
Start your first chapter free →We deleted the instruction that was hurting the reads
Two global instructions were being appended to your text behind the scenes, and both were doing damage.
The first told the engine to use a natural speaking rhythm on every line. In our listening sessions it slowed undirected narration by roughly a third and introduced pauses in the middle of sentences where no pause belonged. The second told the engine to pause naturally on any segment over 200 characters — a length-based rule standing in for a judgment about the sentence's actual structure.
Both are gone. Long sentences are now handled where the problem actually is: run-on and oversized segments are split at their punctuation, so the break falls at a clause boundary the author already wrote, rather than wherever a generic instruction lands it. Targeted pacing repair still runs when our quality check identifies a specific line that needs it — the difference is that it's now specific, applied to a line that has a problem, instead of applied to everything.
Acting directions are now the exception
This is the change underneath the other two. Our segmentation now defaults every line to no direction at all, with one narrow exception: a delivery cue that appears in the text after the line — the trailing "she said, barely above a whisper" that backward-looking context cannot see, because it hasn't been read yet.
That exception is deliberate and small. Everything else the reader would infer from the scene, the engine can now infer too, because the scene is attached.
The measurement, from the first chapter produced entirely under the new engine: 7 of 85 lines carried a direction. Under the previous prompt, roughly 77% did. Of those 7, 6 were exactly the trailing-cue case the rule exists for.
| Previous engine | Current engine | |
|---|---|---|
| Lines carrying an acting direction | ~77% | 7 of 85 (~8%) |
| Context sent with each line | none | up to 5 preceding lines |
| Global pacing instruction | appended to every line | removed |
| Long-sentence handling | generic "pause naturally" | split at punctuation |
Fewer directions is not a cost-saving measure and it is not us doing less work. It is the pipeline trusting the manuscript. A line delivered because of where it sits in a scene reads more like performance than a line delivered because a label told it to sound worried.
One more fix: abbreviations that read in context
A small change with a disproportionate effect on nonfiction and contemporary settings. Abbreviation expansion now happens with awareness of the surrounding text, so "Elm St." resolves to "Elm Street" rather than "Elm Saint" — the kind of error that is invisible in a script and unmissable in a finished chapter. Your custom pronunciations remain authoritative and are unaffected.
What you should do with this
Nothing, if you have work in flight — the engine is live and new productions use it automatically. If you have a chapter you produced before this week and remember it reading a shade slow or a shade flat, that is worth a re-render now.
The best way to judge it is by ear. Listen to a scene with real back-and-forth in it on our published productions — the dialogue in Frankenstein is a fair test — and pay attention to what happens between lines rather than within them.
If you want to shape a specific moment further, that is what the agentic director is for: describe the change in plain language and it applies it to the exact line. And if you want the vocabulary for what you're listening for, our guide to the tells listeners hear in AI narration names the artifacts this release was built to remove.
FAQ
Do I need to re-render existing chapters to get this? Yes. Chapters produced before this release were rendered under the previous engine. New productions and any re-render pick up the new behavior automatically.
Will fewer acting directions make my narration flatter? That was the concern going in, and our listening sessions found the opposite: lines rendered bare with scene context beat directed lines by ear. Direction was compensating for missing context. With the context supplied, the compensation is what was flattening things.
Can I still direct a specific line myself? Yes. Per-line direction is still available, and the agentic director applies it conversationally. The change is to the default, not to your control.
Does the added context cost me more credits? No. Pricing is unchanged — credits are charged per word on your manuscript, and the surrounding-line context is not billed.




