Your audiobook is finished, you upload it, and a machine rejects it in under a minute. No one listened. The rejection notice names a number — RMS, peak, noise floor — and you are left guessing which knob to turn.
The audiobook audio file requirements that trigger those rejections are narrower than most first-time publishers expect, and they are almost identical across distributors. This is the spec sheet: the six numbers every platform checks, the file-structure rules that sit alongside them, and what to do when your master misses.
The Six Numbers
<span id="tldr"></span>
Every mainstream audiobook distribution path runs an automated audio check before a human reviewer sees anything. These are the thresholds, published verbatim by ACX and matched by Author's Republic:
| Requirement | Spec | What it actually measures |
|---|---|---|
| RMS loudness | -23 dBFS to -18 dBFS RMS | Average perceived volume across the whole file |
| Peak level | no higher than -3 dB | The single loudest instant, kept clear of clipping |
| Noise floor | no higher than -60 dB | Room hiss, hum, and preamp noise in the silences |
| Sample rate | 44.1 kHz | Recording and export setting |
| Bitrate | 192 kbps CBR | Constant bitrate, not variable |
| Format | MP3 | Mono or stereo — but the same choice in every file |
Miss any one and the file bounces. There is no "close enough" band: a noise floor at -58 dB fails, and a peak at -2.9 dB fails.
Ready to try it on your own book?
Start your first chapter free →dBFS RMS Is Not LUFS
This is the single most expensive misunderstanding in audiobook mastering, so it is worth taking before anything else.
There are two incompatible mastering conventions in spoken-word audio, and they use different meters:
- Podcast and streaming target LUFS — roughly -16 LUFS integrated for a stereo podcast, true peak no higher than -1 dBTP, usually at 48 kHz.
- Audiobook retail targets the older dBFS RMS convention — -23 to -18 dBFS RMS, peak no higher than -3 dB, at 44.1 kHz.
LUFS is loudness-weighted and dBFS RMS is not, so the two numbers never line up cleanly. One file cannot satisfy both specs. A master built to podcast convention will typically fail an audiobook check on sample rate (48 kHz where 44.1 is required) and on peak (-1 dBTP where -3 dB is the ceiling) — and, counter-intuitively, not on loudness. We measured one of our own older chapter files: 48 kHz, -16.3 LUFS integrated, true peak -1.5 dBFS. It failed on sample rate and peak. The same file reads -19.5 dBFS RMS, sitting comfortably inside the -23 to -18 retail window.
That is why "it sounds right on Spotify" tells you nothing about whether it will clear a distributor. If you produce for both, you need two exports from the same mix, not one file with a compromise setting.
Why the ACX Numbers Are the Industry Numbers
ACX wrote this spec for Audible, and the rest of the market adopted it wholesale. Author's Republic and PublishDrive publish the same RMS, peak, sample-rate, and bitrate values. The Spotify-facing pipeline uses the same thresholds, delivered per chapter.
Worth knowing about that pipeline, because the names changed recently: Findaway Voices was reacquired from Spotify by its original founding team in August 2025 and now operates as Voices by INaudio, while Spotify for Authors handles direct-to-Spotify publishing and analytics. Different brands, same audio spec.
Where distributors differ is at the edges — file length, file size, naming — and those edges are what actually catch authors out, because most indie authors reach the stores through an aggregator rather than uploading to ACX directly:
| Route | Audio | The rule that catches people |
|---|---|---|
| ACX | 192 kbps CBR MP3 or higher, 44.1 kHz, -23 to -18 dBFS RMS, peak ≤ -3 dB, noise floor ≤ -60 dB | 120 minutes per file; separate opening and closing credits files; a 1–5 minute retail sample |
| Author's Republic | ACX-equivalent | A hard 170 MB per file ceiling on top of the duration limit |
| PublishDrive | MP3 or M4B, 192 kbps CBR or higher, 44.1 kHz, -23 to -18 dBFS RMS, peak ≤ -3 dB | 78 minutes per file for Libby (119 for Audible), and every file in the title must share one bitrate and sample rate |
| Voices by INaudio | 192 kbps CBR MP3 or FLAC, 44.1 kHz, -18 dB RMS, peak ≤ -3 dB, noise floor ≤ -60 dB | CBR only — a variable-bitrate export is rejected |
| StreetLib | 192 kbps MP3, 44.1 kHz, 16-bit, RMS -24 to -14 (looser than ACX) | Filenames numbered 001_, 50-character limit, no accented characters |
| Google Play Books | MP3, AAC, FLAC, WAV or ZIP; 44.1 kHz or higher; mono ≥ 128 kbps, stereo ≥ 256 kbps; 5 minutes to 100 hours total | Strict ID_NofM.mp3 naming (9781234567897_1of8.mp3) — wrong names cost about 24 hours of processing delay. Google publishes no RMS, peak, or noise-floor target |
The practical consequence: design around 78 minutes per file, not ACX's 120. A 90-minute chapter clears ACX and then bounces at the distributor most authors actually use.
One consequence matters for anyone producing with AI: ACX does not accept third-party AI or text-to-speech narration — its submission requirements prohibit it outright. So mastering to the ACX spec is not about getting into that particular store. It is about producing files that pass everywhere the spec has been copied, which is nearly everywhere else.
Apple Books has no meaningful direct indie upload path; titles reach it through ACX or an aggregator, each applying its own check. Master to the table above and you satisfy the distributor, which is the gate that matters.
File Structure: the Rules That Are Not About Sound
Roughly half of all rejections have nothing to do with loudness. They are structural:
- One chapter or section per file. No combined files, no split chapters.
- Maximum length per file. ACX allows 120 minutes, but the number to build around is 78 — PublishDrive's Libby limit, and the tightest cap you are likely to meet. Author's Republic adds a hard 170 MB ceiling on top.
- Room tone at both ends. 0.5 to 1 second at the head, 1 to 5 seconds at the tail — silence recorded in the room, not digital zero.
- Consistent channel format. All mono or all stereo. Mixing them across a title fails the check. Full-cast productions with music and effects are usually delivered in stereo, which the spec allows. Every file also has to share one bitrate and sample rate.
- Separate opening and closing credits files. Opening credits state the title, author, and narrator. Closing credits follow the "You have been listening to…" convention and end the book. These are their own files, not the head and tail of chapter one.
- A retail sample. One to five minutes, representative, no spoilers, no explicit content. ACX prefers the first five minutes.
- Clean filenames. Standard alphanumeric characters, with the chapter or section header in the name. Google Play is the exception and wants
ID_NofM.mp3exactly; StreetLib wants a numbered001_prefix under 50 characters with no accents.
The credits and sample requirements catch experienced authors surprisingly often, because they are the parts nobody records until the last day.
What to Do When Your Master Misses
Each failure has a specific fix, and the order matters — noise first, dynamics second, level last.
Noise floor too high. No amount of level adjustment fixes this; normalizing a noisy file raises the noise with the voice. Fix it at the source (quieter room, better mic placement, tighter gain staging) or with a noise-reduction pass applied gently before anything else. Aggressive reduction leaves an artificial, watery texture that a human reviewer will flag even if the meter passes.
RMS out of range. This is a compression and normalization problem. Compress moderately to tighten the dynamic range, then normalize the RMS into the -23 to -18 dBFS window. Chasing RMS with raw gain alone tends to push peaks past -3 dB. Check the meter you are reading: an editor showing LUFS is answering a different question.
Peaks above -3 dB. Limit rather than turn everything down — a brick-wall limiter set a little under the ceiling, around -4 dB, catches plosives and sudden loud lines without flattening the performance and leaves headroom for the encoder.
Wrong sample rate or bitrate. Export settings, not mixing, and the most common single cause of a failed check. Record at 44.1 kHz rather than resampling down from 48 kHz, and export MP3 at 192 kbps constant bitrate. Variable bitrate fails even when the average is higher. If your mix was built for podcast delivery at 48 kHz, re-export it at 44.1 rather than shipping the podcast file.
If you are running a multi-voice production with music and sound design, these steps compound: every element has to sit inside one loudness envelope. That is the part hand-mastering gets expensive, because a full-cast chapter has a dozen sources rather than one voice track.
Where Production Tools Fit
The spec is a mastering problem, and mastering is either your time or someone's invoice. A studio bakes it into the per-finished-hour rate. A DIY workflow makes it your afternoon, per chapter, forever.
Midsummerr produces the mix as part of the production — cast voices, score, and effects balanced into a single chapter master rather than handed to you as raw stems to level yourself. Because retail and podcast are two different specs, the download step asks where the files are going: audiobook retailers, Google Play Books, or a podcast feed. The retail master is 44.1 kHz stereo, 192 kbps CBR, around -20 dBFS RMS with peaks held below the -3 dB ceiling; the podcast option hands you the 48 kHz mix as generated. Files carry room tone at both ends, opening and closing credits are produced as their own files, chapter numbering skips them so "Chapter 1" is the first real chapter, and Google's ID_NofM naming is applied when you pick that destination.
Each file is measured on the way out and the RMS, peak, and noise floor are shown per chapter — the same three numbers Audacity's ACX Check reads. What that does not do is guarantee acceptance. A retailer still reviews by ear, and music under narration can draw a rejection no automated measurement predicts.
Hear what a mixed full-cast chapter sounds like in The Mystery of the Blue Train or Anne of Green Gables. Pricing runs $1.50 per 1,000 words for single-narrator production and $5 per 1,000 words for full production with cast, music, and effects — see pricing for the full breakdown.
Whichever route you take, verify before you upload. Measuring RMS, peak, and noise floor on one chapter takes two minutes in any editor and saves a rejection cycle that can cost a week.
FAQ
What are the audiobook audio file requirements for 2026?
RMS loudness between -23 dBFS and -18 dBFS, peak level no higher than -3 dB, noise floor no higher than -60 dB, 44.1 kHz sample rate, 192 kbps constant-bitrate MP3, mono or stereo used consistently, and one chapter per file with 0.5–1 second of room tone at the head and 1–5 seconds at the tail. ACX allows 120 minutes per file, but distributors cap lower — PublishDrive at 78 minutes — so 78 is the number to build around.
Is -16 LUFS the right loudness for an audiobook?
No. -16 LUFS is the podcast and streaming convention. Audiobook retail uses dBFS RMS, and asks for -23 to -18 dBFS RMS with peaks under -3 dB at 44.1 kHz. The two are different scales measured with different meters, and one file cannot satisfy both — a podcast master usually fails an audiobook check on sample rate and peak rather than on loudness.
Why does my audiobook keep getting rejected?
Almost always one of three things: a 48 kHz export where 44.1 kHz is required, peaks above -3 dB left over from a podcast-style master, or a structural miss — a variable-bitrate export, mixed mono and stereo files, a file over the distributor's duration cap, or missing opening and closing credits. A high noise floor is the classic fourth cause on home-recorded narration.
Do all audiobook distributors use the same specs?
The audio numbers, effectively yes — ACX defined them and Author's Republic, PublishDrive, and Voices by INaudio publish the same values (INaudio states -18 dB RMS rather than a range). The file rules differ: Author's Republic adds a 170 MB per-file cap, PublishDrive caps at 78 minutes for Libby, StreetLib runs a looser -24 to -14 RMS window with strict filename rules, and Google Play Books publishes no loudness target at all but rejects anything not named ID_NofM.mp3.
How long can one audiobook file be?
ACX allows 120 minutes, but that is the wrong number to plan around unless you are uploading to ACX directly. PublishDrive caps files at 78 minutes for Libby and 119 for Audible, and Author's Republic applies a 170 MB size ceiling. Split at 78 minutes and every route accepts the file.
Is Findaway Voices still around?
It operates as Voices by INaudio. Its original founding team reacquired the platform from Spotify in August 2025, and Spotify for Authors now handles direct Spotify publishing and analytics separately. The audio specifications did not change.
Can I submit an AI-narrated audiobook to ACX?
No. ACX's submission requirements prohibit unauthorized text-to-speech, AI, and automated recordings. Other distribution paths do accept AI-narrated titles — see our guide to where to distribute an AI audiobook in 2026 and our breakdown of whether Audible accepts AI narration.
Should my audiobook be mono or stereo?
Either is accepted, as long as every file in the title matches. Single-narrator recordings are commonly delivered in mono. Full-cast productions with music and sound design are usually stereo, where the extra channel carries the score and effects placement.
The audio spec is the least creative part of publishing an audiobook and one of the most common places a finished production stalls. Six numbers, a handful of file rules, and a two-minute check before upload. If you would rather not run that check by hand every chapter, listen to a produced sample and see what comes out the other side already mixed.




