Two pieces of listener research landed in the same season and appear to disagree. The Audio Publishers Association's 2026 consumer survey says stated willingness to try an AI-narrated audiobook fell from 70% to 61% in a year. A blinded test by the same research firm, whose key findings were published on September 11, says listeners rated an AI multi-cast production above a professional human narrator reading the same book. Most of the coverage picked a side.
The right reading is that both are true, and that the test was not measuring what its headline says it measured. It was measuring the cast.
What the blinded test actually compared
Edison Research at SSRS ran the study, and the AI audiobook company Spoken commissioned it; the version listeners heard was Spoken's own production. That is disclosed in the original release and it belongs at the top of any honest write-up. The design, though, is sound: 1,005 US adults who listen to fiction audiobooks were randomly assigned an excerpt of the same sci-fi thriller, either as a professional single-narrator recording or as an unedited AI multi-cast version, and rated it before anyone told them how it was made.
| Measure (blind) | AI multi-cast | Human single narrator |
|---|---|---|
| Favourability | 61% | 53% |
| Perceived narration quality | 66% | 60% |
| Overall engagement | 58% | 49% |
| Purchase intent | 46% | 49% |
| Believed the narration was human | 61% | 65% |
Source: Edison Research at SSRS, released July 14, 2026; webinar key findings September 11, 2026. Purchase intent is described as statistically comparable.
Notice what the comparison changes. It does not swap a human voice for a synthetic one and hold everything else constant. It swaps one narrator for a full cast and swaps human for synthetic in the same move. Two variables, one result. On its own, that table cannot tell you which of the two the listeners were responding to.
Ready to try it on your own book?
Start your first chapter free →The one result that separates the variables
The release contains the sentence that resolves it, and almost nobody quoted it: "For exposition without multiple characters, the human narrator was rated higher."
Read that with the table above. Where the excerpt had characters, the cast won. Where the excerpt was a narrator describing a room, the human won. The preference tracks the presence of characters to voice, not the presence of a machine. Listeners were not saying they prefer AI. They were saying they prefer being cast to, and they said it even when the cast was, by the study's own description, unedited.
The webinar findings say the same thing from the demand side. Among frequent listeners, 81% are interested in distinct voices for each character and 51% are very interested. Asked what drives them to a title, they ranked narration quality and immersion above cost and above celebrity narrators. Edison's Megan Lazovick summarised the study as a signal that quality matters however the narration is produced, which is a careful way of saying the production is what got rated.
Why the label is losing while the format wins
Now put the APA numbers back. Willingness to try "an AI-narrated audiobook" fell nine points to 61%. Only 16% of audiobook listeners have ever heard one. AI titles were 0.03% of 2025 sales revenue. The APA's executive director framed the year as a decline in preference for AI narration, and on the survey's own terms he is right.
But the survey asks about a label, and the blinded test measured an experience, and the study reports how far apart those two are: before hearing the excerpt, 31% of participants said they were likely to listen to an AI audiobook. After hearing it, 65% said so. The same people, one excerpt apart, doubled. Sixteen percent of listeners have heard an AI title; the other 84% are answering the APA's question from what they imagine, and what most people imagine is a flat machine reading a book alone. That is a fair thing to reject. It is also not what was tested.
So the two studies are not in tension. Stated appetite for the category name is falling. Blind preference for a cast production is rising. An author who plans around the first number is planning around a word.
What this changes for an author
Three practical conclusions, none of which require believing the vendor.
The format decision is separable from the technology decision. Whether your book benefits from a cast is a property of the manuscript: how much of the page is dialogue, how many people hold a scene at once, how hard your attributions are working. We laid out how to read your own book for that, and the answer is sometimes no. A quiet first-person memoir is exposition without multiple characters, and this study says a single narrator is the right call for it. A romantasy with four people in every dinner scene is the other case.
Cast size follows scene structure, not budget. The old reason to ration voices was that a human ensemble costs two to three times a single narrator. That constraint is what made the format rare enough that 84% of listeners have never heard one. The right count is set by how many characters share a scene, and a full production on Midsummerr runs $5 per thousand words, cast, score and sound design included, so a 90,000-word novel is about $450 and generates in hours rather than months.
Finish it. The unedited cast in this test still beat a professional. An edited one, with a score that follows the arc and sound design that places each scene, is the version listeners actually finish at higher rates. And finishing is where the label stops mattering: nobody who reached the last chapter of a book goes back to check the survey question. If you want to hear what a finished cast production sounds like, The Murder of Roger Ackroyd is free to stream, and the dinner-party chapters are exactly the multi-character scenes this study is about.
The honest version
This is one study, commissioned by a company whose product was the tested condition, on one book in one genre, and its strongest numbers are the commissioner's. Treat the 61-to-53 margin as a signal, not a law. The independent figures carry the argument anyway: a nine-point fall in willingness to try the label, a doubling in willingness after hearing the thing, four in five frequent listeners asking for distinct character voices, and a human narrator winning wherever there were no characters to cast.
Those four facts describe a market that has already decided what it wants and has not yet been given a name for it. The name is not "AI narration." It is a cast.
Further reading
- The case for dramatized audio just got a number: Audible's completion-rate finding and why it matters for a series.
- Listeners aren't rejecting AI. They're rejecting unfinished work.: the craft argument that this study now quantifies.
- Audiobook listening statistics 2026: the APA market data this post draws on, with sources.




