A full-cast audiobook gives every character in your novel their own voice, with the narrator as one voice among them. For most of the format's history, making one meant a room of actors, a director and a studio calendar, which is why it belonged to bestsellers. This guide shows how to make a full-cast audiobook of your own novel in seven steps: what each step asks of you, how you know it is done, and what the finished audiobook costs.
Listeners are asking for the format. In a 2026 study of more than 1,000 US fiction audiobook listeners run by Edison Research at SSRS, 81% of frequent listeners said they were interested in hearing distinct character voices, and 51% were very interested. One caveat belongs next to that number: the study was commissioned by Spoken, a company that sells AI audiobook production, so treat it as a signal. We went through the full results and their limits when they were published.
The seven steps at a glance
- Decide whether the book needs a cast. The manuscript answers this, not the budget.
- Prepare the manuscript. Final text, clean chapters, nothing you do not want read aloud.
- Sort your characters into three tiers. Must be distinct, benefits from distinct, stays with the narrator.
- Cast for contrast, then audition in pairs. Voices are judged against each other, never alone.
- Settle the names before chapter one. A wrong name repeats through every scene it is in.
- Produce one hard scene, then the book. Listen before you commit.
- Review, edit and export. Fix what your ear catches, then master for the stores you have chosen.
Ready to try it on your own book?
Start your first chapter free →Two ways to do each step
The steps are the same whether actors record your book in a studio or you produce it on Midsummerr. Who does the work at each step is what differs.
| Step | Studio with human actors | On Midsummerr |
|---|---|---|
| Who speaks each line | The script is marked up for every actor | Dialogue is assigned to its speaker automatically |
| Casting | Auditions and a contract for each actor | Every character arrives with a voice and a written description you can change |
| Performance | Directed recording sessions | A directed performance, generated chapter by chapter |
| Music and sound effects | Composed or licensed, then mixed, as separate work | Score and sound design created and placed scene by scene (Full Production) |
| Changing a line afterwards | A pickup session with the actor | Regenerate the line; edits are free |
| What the price follows | Finished hours of audio, and every actor is paid | Words in the manuscript; cast size does not change it |
| Calendar | Set by recording and editing schedules | A few hours to generate, then your own review |
The seven steps in detail
Each step below ends with a test, so you know when to move on.
Step 1: Decide whether the book needs a cast
Not every novel does. Open a scene you already know is difficult and read it for dialogue: how much of the page is speech, how many people speak in one scene, and how hard your "he said, she said" is working to keep them apart. A crowded thriller interrogation needs distinct voices to stay legible. A first-person literary novel that lives inside one head is often stronger with a single narrator.
The Edison study points the same way. Its published findings note that the human single narrator was rated higher for narrator exposition, the passages where nobody is speaking. A cast earns its place in dialogue. Does your book need a full cast? walks through the five signals to read for, and full cast versus single narrator compares what the listener hears.
Done when: you can say in one sentence why this book needs more than one voice.
Step 2: Prepare the manuscript
Whatever is on the page gets performed, so the audiobook starts from the final, proofread text.
- Use the published version. A typo in the manuscript becomes a stumble in the audio.
- Keep dialogue punctuation consistent. Quoted speech is what gets handed to a character.
- Give every chapter a clear heading. Chapters are detected at upload, and you can reorder, split, add or delete them afterwards.
- Cut what should not be read. The table of contents, the copyright page, footnotes and image captions.
On Midsummerr you upload a DOCX, EPUB or TXT file. You do not need to tag who is speaking: dialogue is assigned to the cast for you.
Done when: the file you upload is the text you would be happy to hear word for word.
Step 3: Sort your characters into three tiers
Cast size is set by who shares a scene, not by how many names are in the book.
- Tier 1, must be distinct. Characters who trade lines in dialogue-heavy scenes: your point-of-view characters, the antagonist, and whoever the protagonist argues with most. In most novels this is three to six people.
- Tier 2, benefits from distinct. Recurring characters with real presence who mostly appear alongside a Tier 1 character.
- Tier 3, stays with the narrator. The waiter, the dispatcher, the voice in the crowd.
The common mistake is casting every named character. It flattens the hierarchy your prose already built and makes the audio sound busy. How many character voices your audiobook needs covers the method in full.
Done when: you have a short list of the voices the listener must never confuse.
Step 4: Cast for contrast, then audition in pairs
A voice is only right or wrong next to the voices around it. Two warm, mid-pitched, unhurried leads will blur in the listener's ear however well each one suits its character.
Cast your Tier 1 list along the differences a listener can hear with their eyes closed: pitch, pace, age and accent. Then audition in pairs. Play the samples of the two characters who share the most scenes back to back, and if you have to think about which is which, change one.
On Midsummerr every character arrives with a voice and a written description that opens with the accent. Play each sample, and where a voice is wrong, describe what you want and get a new one. Nothing is contracted, so you can keep going until the pairs separate cleanly. Choosing the right voice for each character goes deeper on matching a voice to the person on the page.
Done when: you can tell every Tier 1 pair apart from the samples alone.
Step 5: Settle the names before chapter one
The mistake listeners forgive least is a mispronounced name, and it is the one authors find last. It sits in chapter one, sounds plausible to anyone who has never met the word, and is repeated in every scene that character is in.
List the names, places and invented words that could be read more than one way and decide each before anything is produced. On English-language productions, Midsummerr flags these at upload, ordered by how often they appear, each with suggested readings you can hear in the narrator's voice. Use one, dismiss it, or type your own. Pronunciation control covers the guide itself.
Done when: you have heard every invented name said the way you say it.
Step 6: Produce one hard scene, then the book
Do not judge a cast on its easiest page. Start with the scene that worried you in Step 1 and listen for five things:
- In the fastest exchange, can you tell who is speaking without the attribution?
- Does the narrator sound like the book?
- Do the two leads stay apart when they are angry, and when they are quiet?
- Is the music under the dialogue, not on top of it?
- Was any name read wrong?
If something is off, go back to Step 4 or 5 and change it now, while it is one scene. Then generate the rest. On Midsummerr the first 5,000 words are free with no card, which is usually enough to hear an opening scene with your own cast, and a full novel takes a few hours to generate.
Done when: you have listened to one scene start to finish and would not change the cast.
Step 7: Review, edit and export
Listen through with the manuscript closed. Your ear will catch a pause that is too short, a line that should land harder, a door that closes too loudly. In the Midsummerr editor you can change a line's delivery, a pause, a music cue or an effect, or tell Martin, the in-editor director, what you want in plain language. Edits are free.
Then export for the stores you have chosen. ACX publishes its numbers in its audio submission requirements: volume between -23 dB and -18 dB RMS, peaks below -3 dB, a noise floor below -60 dB RMS, in a constant-bit-rate MP3 of 192 kbps or higher at 44.1 kHz. Other distributors publish their own, and our spec sheet sets them side by side. Midsummerr's download centre gives you retail masters and shows the measured loudness, peak and noise floor of what you are downloading.
One rule to read before you produce, not after: the same ACX page, last updated on April 15, 2026, states that "unauthorized use of text-to-speech, AI, or automated recordings in ACX titles is prohibited." A digitally voiced full cast goes out through the stores and distributors that accept digitally narrated titles, each with its own disclosure rule. Where to distribute your audiobook is kept current on which routes are open.
Done when: the files are downloaded and you have read the current policy of every store you are submitting to.
What a full-cast audiobook costs
Human audiobook work is priced per finished hour. ACX's budgeting guidance says about 9,300 words make one finished hour and puts a retail-ready production at $300 to $400 per finished hour, roughly $200 for narration and $200 for post-production. A 90,000-word novel is about 9.7 finished hours, so one human narrator costs roughly $2,900 to $3,900. That figure is for a single voice. A human full cast adds a fee for each actor on top.
Midsummerr prices by the word, and the cast size does not enter into it.
| Route for a 90,000-word novel | Cost | What the price follows |
|---|---|---|
| One human narrator, retail-ready | Roughly $2,900 to $3,900 | ACX guidance: $300 to $400 per finished hour, about 9.7 hours |
| Human full cast in a studio | More than one narrator: every actor is paid | Quoted per project |
| Midsummerr Full Cast | $337.50 | $3.75 per 1,000 words |
| Midsummerr Full Production (cast, score, sound design) | $450 | $5 per 1,000 words |
You pay once, you own the files with full commercial rights, and edits are free. What an audiobook costs to produce has the wider comparison, and how long an audiobook takes sets the timelines side by side.
Hear one before you make one
The fastest way to make the decisions in Steps 3 and 4 is to hear a sorted cast at work. Jane Eyre is a full production of more than 18 hours with a full cast, an original score and sound design. Frankenstein is a harder case: a story told by a narrator inside a narrator, where the frame, Victor and the creature each need to stay distinct for the structure to hold in audio. Listen to a crowded scene in either one with your eyes closed, which is the test your own book has to pass.
FAQ
How many voices does a full-cast audiobook need?
Most novels need three to six voices that must be distinct, plus the narrator. Those are the characters who trade lines in dialogue-heavy scenes. Recurring characters benefit from their own voice, and one-line parts are best left with the narrator. Books with several storylines running at once need more.
Do I need to mark who is speaking in my manuscript?
No. On Midsummerr, dialogue is assigned to its speaker automatically, and the narrator keeps the narration. Upload the manuscript as it was published, with consistent quotation marks and clear chapter headings.
How much does it cost to make a full-cast audiobook?
On Midsummerr, Full Cast is $3.75 per 1,000 words and Full Production, which adds an original score and sound design, is $5 per 1,000 words. A 90,000-word novel is $337.50 or $450, and the number of characters does not change the price. The first 5,000 words are free.
How long does it take?
Generating a full novel takes a few hours. What follows is your own review and editing, which takes as long as you want it to. A finished full-cast audiobook in days is the realistic timeline.
Can I change a voice after the audiobook is generated?
Yes. You can direct and regenerate any line, voice, music cue or sound effect, and edits are free. Casting is a decision you can revise after hearing it, which is why Step 6 produces one scene first.
Can I sell a digitally voiced full-cast audiobook on Audible?
Not through ACX, the self-service route to Audible. Its submission requirements prohibit unauthorized text-to-speech, AI or automated recordings. Other stores and distributors accept digitally narrated titles on their own terms, and those terms change: does Audible accept AI narration? and our distribution guide track the current rules.
Start with your hardest scene
Upload your manuscript, meet your cast, and hear your own opening scene on the first 5,000 words for free. If the fastest exchange in the book works with your eyes closed, the rest of the production will too. Start with 5,000 free words or listen to a full production first. For the format itself, our full-cast audiobook guide is the place to begin.




