Sound-designed audiobooks are showing up in more catalogues, and listener reaction is visibly split — an r/audiobooks thread this week ran under the title Wtf is up with these new sound effects?, with the replies falling on both sides. That split is worth taking seriously, because the two camps are usually arguing about different things.
Almost nobody objects to a production having a score. What people object to is a mix. A music bed that competes with a line, an effect that arrives at a volume the narration never reaches, a scene where the words got harder to catch. Those are level decisions, and level decisions are fixable.
Here is the answer up front, before the detail.
<span id="tldr"></span>
Sound effects and music in an audiobook should sit under the voice at every moment, and no effect should ever make a word harder to catch. The voice is the signal. Everything else is context. If a listener has to lean in, rewind, or reach for the volume knob, the mix is wrong — regardless of how the effect sounds on its own.
Why audiobooks are stricter than film
Film mixers work dialogue-first too, but they have a safety net audiobooks do not: a picture.
When a line gets partly buried under a score in a film, the audience still has faces, lips, gestures, and the visible situation. Meaning survives a masked syllable. In an audiobook the words are the entire channel. A listener who misses "he didn't say that" has no second source. They rewind — and rewinding is the moment a production stops being immersive and starts being work.
The second difference is the listening environment. Cinema mixes assume a quiet dark room. Audiobooks are consumed in cars, on trains, in kitchens with an extractor fan on, through one earbud while walking. Every one of those environments has noise that eats the quiet end of your dynamic range. A subtle bed at home is inaudible in a car; a slam that felt cinematic at home is genuinely startling at 70mph.
Design for the worst listening condition your book will meet, not the best one.
Ready to try it on your own book?
Start your first chapter free →The four rules that resolve most complaints
1. Effects live under dialogue, not next to it
The instinct is to place a big effect in the gap between two lines, at full volume, where it has room. That is the loudest possible way to do it, and it is what listeners flag as intrusive.
Prefer effects that run underneath narration at a level where they register as environment rather than as an event. Room tone, weather, a fire, a crowd — these should read as "we are somewhere" and not as "something happened." Save the discrete, foreground effect for the handful of moments in a book that genuinely turn on it.
2. Transients are the real offender
Listeners rarely complain about a string pad. They complain about the door slam, the gunshot, the thunder crack. Sustained sound is easy to mix quietly and still feel; a transient has to be loud to feel like an impact, which is precisely why it hurts.
If a transient is going to hit, it should hit below the loudest narration in the same scene, not above it. The impression of force comes from contrast and from what surrounds it — a beat of near-silence before, a tail after — far more than from raw level.
3. The bed comes down when the voice comes in
Music that works beautifully in a chapter-opening bar is often too present the moment a narrator starts. The fix is not to pick quieter music but to move it: bring the bed up in the gaps, take it down under speech, and let it breathe again at the end of the scene.
A static music level across a chapter is the single most common cause of "the music was too loud" — the level was fine for the intro and wrong for everything after it.
4. Nothing in the mix should require a volume change
This is the practical acceptance test. Play the chapter at one fixed volume, set so the narration is comfortable. If any moment makes you want to turn it down, or any dialogue makes you want to turn it up, the mix is not finished. A listener in a car cannot ride the fader for you.
The delivery specs put a hard ceiling on all of it
Sound design does not get its own budget. It shares the voice's.
Every mainstream distribution path runs an automated audio check on the finished file, and ACX publishes the thresholds that the rest of the market broadly matches:
| Requirement | Spec | What it means for sound design |
|---|---|---|
| RMS loudness | -23 to -18 dBFS RMS | Measured across the whole file. A loud, constant music bed raises the average and can push a correctly-narrated chapter out of the window. |
| Peak level | no higher than -3 dB | Your loudest transient — usually an effect, not a word — owns this number. |
| Noise floor | no higher than -60 dB | Ambient beds and room tone are signal, not noise, but a bed left running under silence can complicate an otherwise clean measurement. |
| Sample rate | 44.1 kHz | Applies to the delivered master regardless of what the elements were made at. |
| Bitrate | 192 kbps CBR or higher | Dense mixes are the ones that suffer most from lossy encoding — check the encoded file, not the source. |
Two consequences follow directly. First, the RMS window applies to the mix, so every decibel of music is a decibel the narration cannot use. Second, peak headroom is almost always spent by an effect, which means your loudest sound effect is functionally setting the ceiling for the entire book.
Note that dBFS RMS is not LUFS, and the two are not interchangeable — we cover that distinction, and the rest of the delivery requirements, in audiobook audio file requirements.
What "well mixed" actually sounds like
The goal is sound design you feel and do not audit. A listener should finish a chapter able to describe the room the scene happened in without being able to tell you what they heard.
Our production of Frankenstein is a reasonable place to test the claim against your own ear: the score and the environments carry the Arctic framing and the laboratory scenes, and the narration stays the thing you are actually following. Play it once at a normal volume and once at a low one. A mix that survives the low pass is a mix that will survive a commute.
If you want the case for why the layer exists at all before you worry about its levels, sound design in audiobooks covers what score and effects do to a story. Why dramatized audiobooks feel immersive takes the listener's side of the same question, and the full-cast audiobook guide covers the casting decisions that sit alongside these mixing ones.
Where this gets decided in practice
In a traditional pipeline, mix decisions are made once, by an engineer, in a session you are not in — and revisiting them means booking more time. That is a large part of why sound design has historically been reserved for lead titles: it is not just the cost of making the effects, it is the cost of changing your mind about them.
On Midsummerr, sound design is part of the production rather than a separate post stage, and the elements on a chapter's timeline stay editable after you hear them. An effect that reads as too loud in context is a thing you adjust and replay, not a thing you re-commission. Full Cast is $3.75 per 1,000 words and Full Production — narration, cast, score, and effects — is $5 per 1,000 words.
The practical difference is not mainly price. It is that "is that slam too loud?" becomes a question you can answer by listening and changing it, in the same afternoon.
FAQ
How loud should sound effects be in an audiobook?
Below the narration at all times, and mixed so no effect makes a word harder to catch. There is no single published number, because it depends on the scene and the narration level — but the test is reliable: if you would reach for the volume knob at any point while the chapter plays at a fixed setting, an effect is too loud.
Do sound effects affect whether my audiobook passes distribution checks?
Yes. The ACX thresholds — -23 to -18 dBFS RMS, peaks no higher than -3 dB — are measured on the finished mix, not on the narration alone. A loud music bed raises average loudness and a loud transient sets your peak, so both can push an otherwise-compliant file out of spec.
Should music play under narration or only between scenes?
Both, at different levels. Beds under narration should be quiet enough to read as atmosphere; music can come up in scene transitions and chapter openings where no words are competing with it. The mistake is leaving one level running across both cases.
Why do some listeners dislike sound effects in audiobooks?
The common complaints are practical: effects loud enough to startle in a car, music that buries dialogue, or sound used to replace descriptive prose rather than support it. Those are production choices rather than properties of the format — which is why the same listeners often have no objection to a well-mixed production.
Is dBFS RMS the same as LUFS?
No. They are different measurements of loudness and cannot be converted reliably between formats. ACX and most audiobook distribution specifies dBFS RMS; a target given in LUFS is a different instruction. See audiobook audio file requirements for the detail.
The short version
Sound design is not what listeners object to — levels are. Keep the voice the loudest thing in the mix, put effects under dialogue rather than between it, watch your transients because they set the ceiling for the whole book, and never ship a chapter that requires the listener to touch the volume. Then check it on a phone.
Hear a full production and judge the balance for yourself.




