Skip to main content
Midsummerr
ListenFeaturesServicesPricingAboutBlog
Sign InGet Started
  1. Blog
  2. /
  3. Guides

Audiobook Audio File Requirements: The 2026 Spec Sheet

Audiobook audio file requirements for 2026: -23 to -18 dBFS RMS, -3 dB peak, -60 dB noise floor, 44.1 kHz, 192 kbps CBR MP3 — plus what each distributor checks, and why a podcast master fails.

Midsummerr|July 22, 2026|11 min read
Watercolor icon of an audio level meter

TL;DR

Almost every audiobook distributor checks the same six numbers: RMS loudness between -23 dBFS and -18 dBFS, peak no higher than -3 dB, noise floor at or below -60 dB, 44.1 kHz sample rate, 192 kbps constant-bitrate MP3, and one chapter per file. Those are dBFS RMS figures, not the LUFS targets podcasts use — a podcast-ready master fails an audiobook check. Files that miss any number are rejected automatically, before a human ever listens.

Hear a full production first

In this article

  1. 01The Six Numbers
  2. 02dBFS RMS Is Not LUFS
  3. 03Why the ACX Numbers Are the Industry Numbers
  4. 04File Structure: the Rules That Are Not About Sound
  5. 05What to Do When Your Master Misses
  6. 06Where Production Tools Fit
  7. 07FAQ

Audio Sample

Hear a production before you read on

A chapter from a published Midsummerr production — full cast, score, and sound design. Judge the format for yourself.

Loading sample...
Open the full listening page

Your audiobook is finished, you upload it, and a machine rejects it in under a minute. No one listened. The rejection notice names a number — RMS, peak, noise floor — and you are left guessing which knob to turn.

The audiobook audio file requirements that trigger those rejections are narrower than most first-time publishers expect, and they are almost identical across distributors. This is the spec sheet: the six numbers every platform checks, the file-structure rules that sit alongside them, and what to do when your master misses.

The Six Numbers

<span id="tldr"></span>

Every mainstream audiobook distribution path runs an automated audio check before a human reviewer sees anything. These are the thresholds, published verbatim by ACX and matched by Author's Republic:

RequirementSpecWhat it actually measures
RMS loudness-23 dBFS to -18 dBFS RMSAverage perceived volume across the whole file
Peak levelno higher than -3 dBThe single loudest instant, kept clear of clipping
Noise floorno higher than -60 dBRoom hiss, hum, and preamp noise in the silences
Sample rate44.1 kHzRecording and export setting
Bitrate192 kbps CBRConstant bitrate, not variable
FormatMP3Mono or stereo — but the same choice in every file

Miss any one and the file bounces. There is no "close enough" band: a noise floor at -58 dB fails, and a peak at -2.9 dB fails.

Ready to try it on your own book?

Start your first chapter free →

dBFS RMS Is Not LUFS

This is the single most expensive misunderstanding in audiobook mastering, so it is worth taking before anything else.

There are two incompatible mastering conventions in spoken-word audio, and they use different meters:

  • Podcast and streaming target LUFS — roughly -16 LUFS integrated for a stereo podcast, true peak no higher than -1 dBTP, usually at 48 kHz.
  • Audiobook retail targets the older dBFS RMS convention — -23 to -18 dBFS RMS, peak no higher than -3 dB, at 44.1 kHz.

LUFS is loudness-weighted and dBFS RMS is not, so the two numbers never line up cleanly. One file cannot satisfy both specs. A master built to podcast convention will typically fail an audiobook check on sample rate (48 kHz where 44.1 is required) and on peak (-1 dBTP where -3 dB is the ceiling) — and, counter-intuitively, not on loudness. We measured one of our own older chapter files: 48 kHz, -16.3 LUFS integrated, true peak -1.5 dBFS. It failed on sample rate and peak. The same file reads -19.5 dBFS RMS, sitting comfortably inside the -23 to -18 retail window.

That is why "it sounds right on Spotify" tells you nothing about whether it will clear a distributor. If you produce for both, you need two exports from the same mix, not one file with a compromise setting.

Why the ACX Numbers Are the Industry Numbers

ACX wrote this spec for Audible, and the rest of the market adopted it wholesale. Author's Republic and PublishDrive publish the same RMS, peak, sample-rate, and bitrate values. The Spotify-facing pipeline uses the same thresholds, delivered per chapter.

Worth knowing about that pipeline, because the names changed recently: Findaway Voices was reacquired from Spotify by its original founding team in August 2025 and now operates as Voices by INaudio, while Spotify for Authors handles direct-to-Spotify publishing and analytics. Different brands, same audio spec.

Where distributors differ is at the edges — file length, file size, naming — and those edges are what actually catch authors out, because most indie authors reach the stores through an aggregator rather than uploading to ACX directly:

RouteAudioThe rule that catches people
ACX192 kbps CBR MP3 or higher, 44.1 kHz, -23 to -18 dBFS RMS, peak ≤ -3 dB, noise floor ≤ -60 dB120 minutes per file; separate opening and closing credits files; a 1–5 minute retail sample
Author's RepublicACX-equivalentA hard 170 MB per file ceiling on top of the duration limit
PublishDriveMP3 or M4B, 192 kbps CBR or higher, 44.1 kHz, -23 to -18 dBFS RMS, peak ≤ -3 dB78 minutes per file for Libby (119 for Audible), and every file in the title must share one bitrate and sample rate
Voices by INaudio192 kbps CBR MP3 or FLAC, 44.1 kHz, -18 dB RMS, peak ≤ -3 dB, noise floor ≤ -60 dBCBR only — a variable-bitrate export is rejected
StreetLib192 kbps MP3, 44.1 kHz, 16-bit, RMS -24 to -14 (looser than ACX)Filenames numbered 001_, 50-character limit, no accented characters
Google Play BooksMP3, AAC, FLAC, WAV or ZIP; 44.1 kHz or higher; mono ≥ 128 kbps, stereo ≥ 256 kbps; 5 minutes to 100 hours totalStrict ID_NofM.mp3 naming (9781234567897_1of8.mp3) — wrong names cost about 24 hours of processing delay. Google publishes no RMS, peak, or noise-floor target

The practical consequence: design around 78 minutes per file, not ACX's 120. A 90-minute chapter clears ACX and then bounces at the distributor most authors actually use.

One consequence matters for anyone producing with AI: ACX does not accept third-party AI or text-to-speech narration — its submission requirements prohibit it outright. So mastering to the ACX spec is not about getting into that particular store. It is about producing files that pass everywhere the spec has been copied, which is nearly everywhere else.

Apple Books has no meaningful direct indie upload path; titles reach it through ACX or an aggregator, each applying its own check. Master to the table above and you satisfy the distributor, which is the gate that matters.

File Structure: the Rules That Are Not About Sound

Roughly half of all rejections have nothing to do with loudness. They are structural:

  • One chapter or section per file. No combined files, no split chapters.
  • Maximum length per file. ACX allows 120 minutes, but the number to build around is 78 — PublishDrive's Libby limit, and the tightest cap you are likely to meet. Author's Republic adds a hard 170 MB ceiling on top.
  • Room tone at both ends. 0.5 to 1 second at the head, 1 to 5 seconds at the tail — silence recorded in the room, not digital zero.
  • Consistent channel format. All mono or all stereo. Mixing them across a title fails the check. Full-cast productions with music and effects are usually delivered in stereo, which the spec allows. Every file also has to share one bitrate and sample rate.
  • Separate opening and closing credits files. Opening credits state the title, author, and narrator. Closing credits follow the "You have been listening to…" convention and end the book. These are their own files, not the head and tail of chapter one.
  • A retail sample. One to five minutes, representative, no spoilers, no explicit content. ACX prefers the first five minutes.
  • Clean filenames. Standard alphanumeric characters, with the chapter or section header in the name. Google Play is the exception and wants ID_NofM.mp3 exactly; StreetLib wants a numbered 001_ prefix under 50 characters with no accents.

The credits and sample requirements catch experienced authors surprisingly often, because they are the parts nobody records until the last day.

What to Do When Your Master Misses

Each failure has a specific fix, and the order matters — noise first, dynamics second, level last.

Noise floor too high. No amount of level adjustment fixes this; normalizing a noisy file raises the noise with the voice. Fix it at the source (quieter room, better mic placement, tighter gain staging) or with a noise-reduction pass applied gently before anything else. Aggressive reduction leaves an artificial, watery texture that a human reviewer will flag even if the meter passes.

RMS out of range. This is a compression and normalization problem. Compress moderately to tighten the dynamic range, then normalize the RMS into the -23 to -18 dBFS window. Chasing RMS with raw gain alone tends to push peaks past -3 dB. Check the meter you are reading: an editor showing LUFS is answering a different question.

Peaks above -3 dB. Limit rather than turn everything down — a brick-wall limiter set a little under the ceiling, around -4 dB, catches plosives and sudden loud lines without flattening the performance and leaves headroom for the encoder.

Wrong sample rate or bitrate. Export settings, not mixing, and the most common single cause of a failed check. Record at 44.1 kHz rather than resampling down from 48 kHz, and export MP3 at 192 kbps constant bitrate. Variable bitrate fails even when the average is higher. If your mix was built for podcast delivery at 48 kHz, re-export it at 44.1 rather than shipping the podcast file.

If you are running a multi-voice production with music and sound design, these steps compound: every element has to sit inside one loudness envelope. That is the part hand-mastering gets expensive, because a full-cast chapter has a dozen sources rather than one voice track.

Where Production Tools Fit

The spec is a mastering problem, and mastering is either your time or someone's invoice. A studio bakes it into the per-finished-hour rate. A DIY workflow makes it your afternoon, per chapter, forever.

Midsummerr produces the mix as part of the production — cast voices, score, and effects balanced into a single chapter master rather than handed to you as raw stems to level yourself. Because retail and podcast are two different specs, the download step asks where the files are going: audiobook retailers, Google Play Books, or a podcast feed. The retail master is 44.1 kHz stereo, 192 kbps CBR, around -20 dBFS RMS with peaks held below the -3 dB ceiling; the podcast option hands you the 48 kHz mix as generated. Files carry room tone at both ends, opening and closing credits are produced as their own files, chapter numbering skips them so "Chapter 1" is the first real chapter, and Google's ID_NofM naming is applied when you pick that destination.

Each file is measured on the way out and the RMS, peak, and noise floor are shown per chapter — the same three numbers Audacity's ACX Check reads. What that does not do is guarantee acceptance. A retailer still reviews by ear, and music under narration can draw a rejection no automated measurement predicts.

Hear what a mixed full-cast chapter sounds like in The Mystery of the Blue Train or Anne of Green Gables. Pricing runs $1.50 per 1,000 words for single-narrator production and $5 per 1,000 words for full production with cast, music, and effects — see pricing for the full breakdown.

Whichever route you take, verify before you upload. Measuring RMS, peak, and noise floor on one chapter takes two minutes in any editor and saves a rejection cycle that can cost a week.

FAQ

What are the audiobook audio file requirements for 2026?

RMS loudness between -23 dBFS and -18 dBFS, peak level no higher than -3 dB, noise floor no higher than -60 dB, 44.1 kHz sample rate, 192 kbps constant-bitrate MP3, mono or stereo used consistently, and one chapter per file with 0.5–1 second of room tone at the head and 1–5 seconds at the tail. ACX allows 120 minutes per file, but distributors cap lower — PublishDrive at 78 minutes — so 78 is the number to build around.

Is -16 LUFS the right loudness for an audiobook?

No. -16 LUFS is the podcast and streaming convention. Audiobook retail uses dBFS RMS, and asks for -23 to -18 dBFS RMS with peaks under -3 dB at 44.1 kHz. The two are different scales measured with different meters, and one file cannot satisfy both — a podcast master usually fails an audiobook check on sample rate and peak rather than on loudness.

Why does my audiobook keep getting rejected?

Almost always one of three things: a 48 kHz export where 44.1 kHz is required, peaks above -3 dB left over from a podcast-style master, or a structural miss — a variable-bitrate export, mixed mono and stereo files, a file over the distributor's duration cap, or missing opening and closing credits. A high noise floor is the classic fourth cause on home-recorded narration.

Do all audiobook distributors use the same specs?

The audio numbers, effectively yes — ACX defined them and Author's Republic, PublishDrive, and Voices by INaudio publish the same values (INaudio states -18 dB RMS rather than a range). The file rules differ: Author's Republic adds a 170 MB per-file cap, PublishDrive caps at 78 minutes for Libby, StreetLib runs a looser -24 to -14 RMS window with strict filename rules, and Google Play Books publishes no loudness target at all but rejects anything not named ID_NofM.mp3.

How long can one audiobook file be?

ACX allows 120 minutes, but that is the wrong number to plan around unless you are uploading to ACX directly. PublishDrive caps files at 78 minutes for Libby and 119 for Audible, and Author's Republic applies a 170 MB size ceiling. Split at 78 minutes and every route accepts the file.

Is Findaway Voices still around?

It operates as Voices by INaudio. Its original founding team reacquired the platform from Spotify in August 2025, and Spotify for Authors now handles direct Spotify publishing and analytics separately. The audio specifications did not change.

Can I submit an AI-narrated audiobook to ACX?

No. ACX's submission requirements prohibit unauthorized text-to-speech, AI, and automated recordings. Other distribution paths do accept AI-narrated titles — see our guide to where to distribute an AI audiobook in 2026 and our breakdown of whether Audible accepts AI narration.

Should my audiobook be mono or stereo?

Either is accepted, as long as every file in the title matches. Single-narrator recordings are commonly delivered in mono. Full-cast productions with music and sound design are usually stereo, where the extra channel carries the score and effects placement.

The audio spec is the least creative part of publishing an audiobook and one of the most common places a finished production stalls. Six numbers, a handful of file rules, and a two-minute check before upload. If you would rather not run that check by hand every chapter, listen to a produced sample and see what comes out the other side already mixed.

Key takeaways

  • The ACX spec is the de facto industry spec — Author's Republic, PublishDrive, and the Spotify/INaudio pipeline check the same numbers.
  • RMS -23 to -18 dBFS, peak ≤ -3 dB, noise floor ≤ -60 dB, 44.1 kHz, 192 kbps CBR MP3, mono or stereo but consistent.
  • dBFS RMS is not LUFS. A -16 LUFS, 48 kHz podcast master fails on sample rate and peak even when its loudness is fine.
  • ACX allows 120 minutes per file, but distributors cap lower — PublishDrive at 78 minutes. Design around 78.
  • Room tone: 0.5–1 second at the head, 1–5 seconds at the tail. Silence recorded in the room, not digital zero.
  • Opening and closing credits and a retail sample are content requirements, not audio ones — and they get files bounced just as often.
  • ACX does not accept third-party AI narration at all, so meeting its spec is about the standard, not about that store.

Ready to hear your own book like this?

Full cast, original score, sound design. Generation runs in hours — what you change after that is your call.

Start your first chapterListen to Examples

Keep reading

Watercolor ribbon microphone with a flowing voice trail for author-narrated audiobook production
GuidesUpdated

Author-Narrated Audiobooks: When Your Voice Is the Right Call

Should you narrate your own audiobook? The real time cost, when an author's voice works, when it doesn't, and how voice cloning changes the math.

July 16, 2026·9 min read
Watercolor open notebook and microphone on a desk
Guides

Nonfiction and Memoir Audiobook Production

Nonfiction and memoir need different production decisions than fiction: what to do with notes, citations and tables, who reads quoted subjects, and where a cast helps.

July 28, 2026·8 min read
Watercolor headphones resting on an open book
Guides

The AI Narrator "Tells": What Listeners Actually Hear

Listeners can usually name the moment an audiobook stopped sounding produced. Here are the six tells they report, what causes each one, and which are fixable before you publish.

July 27, 2026·7 min read
Watercolor megaphone with sound waves
Guides

How to Market Your Audiobook: The Indie Author Playbook (2026)

How to market your audiobook in 2026: launch pricing, Chirp deals, promo codes, short audio clips, podcast crossover, and Meta ads — a practical indie playbook.

July 20, 2026·6 min read

Midsummerr

Create premium audiobooks with cinematic quality in one click

[email protected]

Quick Links

HomeFeaturesServicesPricingAbout Us

Resources

BlogSupportRequest Demo

Legal

Terms of ServicePrivacy PolicyRefund Policy

© 2026 Midsummerr. All rights reserved.