If you are producing a nonfiction or memoir audiobook, the hard problems arrive earlier than you expect — and none of them are about performance.
A novel is written to be spoken. Nonfiction is written to be read: it carries endnote markers, parenthetical citations, tables, figures, block quotes, and URLs that no one can say aloud as written. Memoir carries a different burden — many real people talking in scenes, across decades, filtered through one retrospective voice. Both break a straight read-through in ways a thriller never does.
It is worth noting where these categories sit commercially. In the Audio Publishers Association's 2026 research, US audiobook revenue reached $2.43 billion in 2025, up 9% year over year, with General Fiction the largest single share at 27% — and the top revenue genres it names are all fiction categories (Publishers Weekly). Nonfiction and memoir are not competing in an immersion arms race. They are competing on credibility and on whether a listener stays to the end. That changes what you should spend production effort on.
This guide covers the four decisions that determine whether your book survives the move to audio. For the broader format primer, start with our full-cast audiobook guide.
Decision one: what happens to the apparatus
Every nonfiction manuscript contains material that exists for the eye. Before production, go through it and assign each type an explicit rule. Do it once, in writing, and apply it consistently — inconsistency is what listeners actually notice.
| Manuscript element | Common audio handling |
|---|---|
| Endnote / footnote markers | Dropped from the read; substantive notes either folded into the sentence or read at chapter end |
| Parenthetical citations (Author, 2024) | Dropped, or reduced to a spoken attribution where the source matters to the argument |
| Bibliography, index | Not read; noted as available in the print and ebook editions |
| Tables and data-heavy figures | Rewritten as a spoken summary of what the table shows, not read cell by cell |
| Charts and images | Described in one or two sentences, or cut if the surrounding text already carries the point |
| URLs and long codes | Replaced with a plain-language pointer ("linked in the companion notes") |
| Block quotes | Read, but marked — by a pause, a shift in voice, or a second voice |
| Headers, callouts, sidebars | Kept, but signposted so the listener knows the main thread paused |
Two rules make this manageable. First, if a passage only makes sense with the figure in front of you, it needs rewriting for audio, not reading. Second, whatever you drop, say so once at the top of the book — a single line in the opening credits that points listeners to a companion PDF or a page on your site is standard and removes the confusion.
Memoir has a lighter version of the same job: photo inserts, letters reproduced as images, family trees. A family tree read aloud is a list of names. Rewritten as one sentence of relationships, it works.
Ready to try it on your own book?
Start your first chapter free →Decision two: who reads the people
This is where nonfiction and memoir diverge, and where most of the production value sits.
Memoir usually wants one voice. The form is a single consciousness looking back, and its power comes from the gap between the person who lived it and the person telling it. Handing dialogue to a cast can flatten that gap by making the past as immediate as the present. One narrator performing everyone keeps the memory inside the rememberer's head, which is where memoir puts it. Listen to how a first-person retrospective carries a whole book in Frankenstein — it opens in Walton's letters and then hands the book to Victor's account of his own past, so every other person in it arrives through a remembering narrator. That is the memoir shape, in fiction.
If that voice should be yours, that is a separate decision with its own tradeoffs, and we have covered it in detail in author-narrated audiobooks.
Reported nonfiction is the opposite case. A book built on interviews, oral history, case studies, or court and archival records has a structural problem in audio: the listener cannot see quotation marks. When three sources are quoted in the same section and the attributions are trailing ("she said", "he told me later"), a listener has to hold the ambiguity for a full sentence before the speaker resolves. Do that forty times and comprehension leaks.
Giving quoted subjects their own voices is not dramatization — it is punctuation. The listener knows instantly that this is a quoted person and which one, and your prose no longer has to carry attribution work the audio can do for free. The same applies to a business book's case studies, a history's primary documents, and a science book's researcher quotes.
Two practical constraints:
- Do not cast real, living, identifiable people as performances. A neutral distinct voice for quoted material reads as a citation device. An impression reads as a claim about how that person sounds, and in nonfiction that is a credibility and rights problem, not a craft choice.
- Keep the count low. Two or three quoted-voice slots, reused consistently by source type, work better than a dozen one-off voices. Consistency is the signal; variety is noise.
Decision three: pronunciation is a credibility issue
In fiction, a mispronounced invented word is an irritation. In nonfiction, a mispronounced real one is a verdict.
Nonfiction manuscripts are dense with things that must be right: researchers and authors you cite, institutions, place names, non-English words, medical and legal terms, acronyms that are sometimes spelled out and sometimes said as words. Memoir adds family names, home towns, and the specific local pronunciations that a reader from that place will hear immediately.
The failure mode is not one error. A cited surname appearing across a chapter, or a place name recurring through a whole memoir, gets said dozens of times. Get it wrong and every instance tells the listener that nobody in the production checked — which, in a book whose whole proposition is that you did the work, is expensive.
So build the list before production, not during review:
- Extract every proper noun and technical term — people, institutions, places, non-English words, acronyms, drug and case names.
- Decide the reading, including stress, and note which acronyms are spoken as letters and which as words.
- Hear it before it propagates through the book.
This is the part of the workflow that wants tooling rather than a document. Midsummerr's pronunciation step is built for it: define an entry as a phonetic respelling, IPA, or a phoneme-style prompt — whichever renders the word correctly — test it in real audio, confirm it, and apply it across affected chapters, with any chapter still waiting on the fix flagged so a correction cannot silently miss half the book. The mechanics are in audiobook pronunciation control.
Decision four: how much production the book actually wants
Nonfiction is the category where restraint pays. A history with a scored montage under every chapter opening sounds like a documentary trailer; a self-help book with atmospheric beds under the exercises sounds like a meditation app. What genuinely helps is structural: a short consistent motif that marks a part or chapter boundary, so a listener at half speed in the car knows the argument moved. Memoir tolerates a little more — an era, a place, a recurring sound tied to one memory — but the same rule holds. Sound that marks structure earns its place; sound that decorates prose does not.
That maps onto tier choice. Midsummerr prices per 1,000 words at $1.50 for a single narrator, $3.75 for full cast, and $5 for full production (cast plus score and sound design):
| Book type | Usually the right shape |
|---|---|
| Memoir, personal essay | Single narrator — the intimacy is the format |
| Self-help, business, how-to | Single narrator, with chapter-boundary marking |
| Reported nonfiction, oral history, true crime | Cast for quoted subjects — comprehension, not drama |
| Narrative history, biography with scenes | Cast, plus restrained score at structural boundaries |
| Immersive narrative nonfiction | Full production, if the book already reads cinematically |
If you are weighing this against studio quotes, our audiobook production cost breakdown has the comparison. And whichever tier you pick, the delivered files still have to meet retailer specs — audiobook audio file requirements covers what those are.
The order that works
Do these in sequence and nothing has to be redone:
- Apparatus pass — mark every note, citation, table, figure and URL with its rule, and rewrite what needs rewriting.
- Voice plan — one narrator or a quoted-voice scheme, decided by whether the book is remembered or reported.
- Pronunciation list — every real name and term, tested in audio, confirmed before the full book generates.
- Structure sound — a boundary motif if the book needs one, nothing that merely decorates.
- Generate, review, correct — review with the manuscript open, because in nonfiction the errors are factual, not emotional.
FAQ
Should I narrate my own memoir?
Often yes — for memoir the author's voice is part of the proposition, and listeners come for it. The honest constraint is duration and consistency: a 90,000-word book is roughly nine to eleven finished hours, and that is a real commitment in stamina, not just time. We go through the tradeoffs, and the hybrid options, in author-narrated audiobooks.
What do you do with endnotes and citations in an audiobook?
Marker numbers and parenthetical citations are normally dropped from the spoken text, since they carry no meaning aloud. Notes that contain real argument are either folded into the sentence or read together at the end of the chapter, and the bibliography is left to the print and ebook editions. Whatever you choose, say it once in the opening credits and stay consistent.
Do nonfiction audiobooks need a full cast?
Not usually — but reported nonfiction benefits from distinct voices for quoted subjects, because audio has no quotation marks. Treat it as an attribution device with two or three reused voice slots, not as dramatization, and never as an impression of a real identifiable person.
How do you handle tables and charts in audio?
Summarize what the table shows in one or two spoken sentences rather than reading it cell by cell, and point listeners to a companion PDF for the full data. If a passage is unintelligible without the figure, rewrite the passage for audio.
Is memoir a good fit for dramatized production?
Memoir generally works better with restraint than with a full dramatization, because its effect depends on a single remembering voice. Narrative history and immersive narrative nonfiction are the nonfiction shapes where cast and score add the most.
What to do next
Nonfiction and memoir do not reward the biggest production. They reward the one that gets the apparatus right, keeps quoted people unambiguous, says every real name correctly, and then stays out of the way.
Hear how a single retrospective voice carries a full-length book in Frankenstein, or browse the rest of our listening page, then bring in your manuscript and start with the pronunciation list — it is the step that saves the most rework.




