Skip to main content
Midsummerr
ListenFeaturesServicesPricingAboutBlog
Sign InGet Started
  1. Blog
  2. /
  3. Guides

Audiobook Audio File Requirements: The 2026 Spec Sheet

Audiobook audio file requirements for 2026: -23 to -18 dB RMS, -3 dB peak, -60 dB noise floor, 44.1 kHz, 192 kbps CBR MP3 — plus what each distributor checks.

Midsummerr|July 22, 2026|6 min read
Watercolor icon of an audio level meter

TL;DR

Almost every audiobook distributor checks the same six numbers: RMS loudness between -23 dB and -18 dB, peak no higher than -3 dB, noise floor at or below -60 dB RMS, 44.1 kHz sample rate, 192 kbps constant-bitrate MP3, and one chapter per file under two hours. Files that miss any of them are rejected automatically, before a human ever listens.

Ready to price your audiobook? Compare Self-Serve, Director-Led, and Voice Conversion →

In this article

  1. 01The Six Numbers
  2. 02Why the ACX Numbers Are the Industry Numbers
  3. 03File Structure: the Rules That Are Not About Sound
  4. 04What to Do When Your Master Misses
  5. 05Where Production Tools Fit
  6. 06FAQ

Your audiobook is finished, you upload it, and a machine rejects it in under a minute. No one listened. The rejection notice names a number — RMS, peak, noise floor — and you are left guessing which knob to turn.

The audiobook audio file requirements that trigger those rejections are narrower than most first-time publishers expect, and they are almost identical across distributors. This is the spec sheet: the six numbers every platform checks, the file-structure rules that sit alongside them, and what to do when your master misses.

The Six Numbers

<span id="tldr"></span>

Every mainstream audiobook distribution path runs an automated audio check before a human reviewer sees anything. These are the thresholds, published verbatim by ACX and matched by Author's Republic:

RequirementSpecWhat it actually measures
RMS loudness-23 dB to -18 dB RMSAverage perceived volume across the whole file
Peak levelno higher than -3 dBThe single loudest instant, kept clear of clipping
Noise floorno higher than -60 dB RMSRoom hiss, hum, and preamp noise in the silences
Sample rate44.1 kHzRecording and export setting
Bitrate192 kbps CBRConstant bitrate, not variable
FormatMP3Mono or stereo — but the same choice in every file

Miss any one and the file bounces. There is no "close enough" band: a noise floor at -58 dB fails, and a peak at -2.9 dB fails.

Ready to try it yourself?

Create your first audiobook free →

Why the ACX Numbers Are the Industry Numbers

ACX wrote this spec for Audible, and the rest of the market adopted it wholesale. Author's Republic publishes the same RMS, peak, noise-floor, sample-rate, and bitrate values. The Spotify-facing pipeline uses the same thresholds, delivered per chapter.

Worth knowing about that pipeline, because the names changed recently: Findaway Voices was reacquired from Spotify by its original founding team in August 2025 and now operates as Voices by INaudio, while Spotify for Authors handles direct-to-Spotify publishing and analytics. Different brands, same audio spec.

One consequence matters for anyone producing with AI: ACX does not accept third-party AI or text-to-speech narration — its submission requirements prohibit it outright. So mastering to the ACX spec is not about getting into that particular store. It is about producing files that pass everywhere the spec has been copied, which is nearly everywhere else.

Apple Books does not publish a comparable numeric spec sheet publicly; titles reach it through distributors who apply their own checks. Master to the table above and you satisfy the distributor, which is the gate that matters.

File Structure: the Rules That Are Not About Sound

Roughly half of all rejections have nothing to do with loudness. They are structural:

  • One chapter or section per file. No combined files, no split chapters.
  • Maximum length per file. ACX caps files at 120 minutes; Author's Republic specifies 119 minutes and a 170 MB ceiling. Stay under both and you are safe everywhere.
  • 1–5 seconds of room tone at the head and tail of every file — silence recorded in the room, not digital zero.
  • Consistent channel format. All mono or all stereo. Mixing them across a title fails the check. Full-cast productions with music and effects are usually delivered in stereo, which the spec allows.
  • Opening and closing credits. Opening credits state the title, author, and narrator. Closing credits follow the "You have been listening to…" convention and end the book. Author's Republic caps each at two minutes.
  • A retail sample. One to five minutes, representative, no spoilers, no explicit content. ACX prefers the first five minutes.
  • Clean filenames. Standard alphanumeric characters, with the chapter or section header in the name.

The credits and sample requirements catch experienced authors surprisingly often, because they are the parts nobody records until the last day.

What to Do When Your Master Misses

Each failure has a specific fix, and the order matters — noise first, dynamics second, level last.

Noise floor too high. No amount of level adjustment fixes this; normalizing a noisy file raises the noise with the voice. Fix it at the source (quieter room, better mic placement, tighter gain staging) or with a noise-reduction pass applied gently before anything else. Aggressive reduction leaves an artificial, watery texture that a human reviewer will flag even if the meter passes.

RMS out of range. This is a compression and normalization problem. Compress moderately to tighten the dynamic range, then normalize the RMS into the -23 to -18 dB window. Chasing RMS with raw gain alone tends to push peaks past -3 dB.

Peaks above -3 dB. Limit rather than turn everything down — a brick-wall limiter at -3.5 dB catches plosives and sudden loud lines without flattening the performance.

Wrong sample rate or bitrate. Export settings, not mixing. Record at 44.1 kHz rather than resampling down from 48 kHz, and export MP3 at 192 kbps constant bitrate. Variable bitrate fails the check even when the average is higher.

If you are running a multi-voice production with music and sound design, these steps compound: every element has to sit inside one loudness envelope. That is the part hand-mastering gets expensive, because a full-cast chapter has a dozen sources rather than one voice track.

Where Production Tools Fit

The spec is a mastering problem, and mastering is either your time or someone's invoice. A studio bakes it into the per-finished-hour rate. A DIY workflow makes it your afternoon, per chapter, forever.

Midsummerr produces the mix as part of the production — cast voices, score, and effects balanced into a single chapter master rather than handed to you as raw stems to level yourself. Hear what a mixed full-cast chapter sounds like in The Mystery of the Blue Train or Anne of Green Gables. Pricing runs $1.50 per 1,000 words for single-narrator production and $5 per 1,000 words for full production with cast, music, and effects — see pricing for the full breakdown.

Whichever route you take, verify before you upload. Measuring RMS, peak, and noise floor on one chapter takes two minutes in any editor and saves a rejection cycle that can cost a week.

FAQ

What are the audiobook audio file requirements for 2026?

RMS loudness between -23 dB and -18 dB, peak level no higher than -3 dB, noise floor no higher than -60 dB RMS, 44.1 kHz sample rate, 192 kbps constant-bitrate MP3, mono or stereo used consistently, and one chapter per file under 120 minutes with 1–5 seconds of room tone at each end.

Why does my audiobook keep getting rejected?

Almost always one of three things: a noise floor above -60 dB from a room that is not quiet enough, RMS outside the -23 to -18 dB window, or a structural miss — a variable-bitrate export, mixed mono and stereo files, or missing opening and closing credits.

Do all audiobook distributors use the same specs?

Effectively yes. ACX defined the spec and Author's Republic and the Spotify/INaudio pipeline publish the same numbers. Apple Books does not publish a public numeric spec and is reached through distributors, who apply their own checks against the same standard.

Is Findaway Voices still around?

It operates as Voices by INaudio. Its original founding team reacquired the platform from Spotify in August 2025, and Spotify for Authors now handles direct Spotify publishing and analytics separately. The audio specifications did not change.

Can I submit an AI-narrated audiobook to ACX?

No. ACX's submission requirements prohibit unauthorized text-to-speech, AI, and automated recordings. Other distribution paths do accept AI-narrated titles — see our guide to where to distribute an AI audiobook in 2026 and our breakdown of whether Audible accepts AI narration.

Should my audiobook be mono or stereo?

Either is accepted, as long as every file in the title matches. Single-narrator recordings are commonly delivered in mono. Full-cast productions with music and sound design are usually stereo, where the extra channel carries the score and effects placement.

The audio spec is the least creative part of publishing an audiobook and one of the most common places a finished production stalls. Six numbers, a handful of file rules, and a two-minute check before upload. If you would rather not run that check by hand every chapter, listen to a produced sample and see what comes out the other side already mixed.

Key takeaways

  • The ACX spec is the de facto industry spec — Author's Republic and the Spotify/INaudio pipeline check the same numbers.
  • RMS -23 to -18 dB, peak ≤ -3 dB, noise floor ≤ -60 dB RMS, 44.1 kHz, 192 kbps CBR MP3, mono or stereo but consistent.
  • One chapter per file, under 120 minutes, with 1–5 seconds of room tone at the head and tail.
  • Opening and closing credits and a retail sample are content requirements, not audio ones — and they get files bounced just as often.
  • ACX does not accept third-party AI narration at all, so meeting its spec is about the standard, not about that store.

Ready to turn your book into a cinematic audiobook?

Full-cast AI voices, original music, and sound effects — production-ready in hours, not months.

Get Started FreeListen to Examples

Keep reading

Watercolor ribbon microphone with a flowing voice trail for author-narrated audiobook production
GuidesUpdated

Author-Narrated Audiobooks: When Your Voice Is the Right Call

Should you narrate your own audiobook? The real time cost, when an author's voice works, when it doesn't, and how voice cloning changes the math.

July 16, 2026·9 min read
Watercolor megaphone with sound waves
Guides

How to Market Your Audiobook: The Indie Author Playbook (2026)

How to market your audiobook in 2026: launch pricing, Chirp deals, promo codes, short audio clips, podcast crossover, and Meta ads — a practical indie playbook.

July 20, 2026·6 min read
Generated watercolor icon of a metronome representing audiobook pacing
Guides

Audiobook Pacing: How Pauses and Silence Shape a Performance

How pacing and pauses work in audiobook production — why silence carries meaning, how rhythm changes by genre, and how to control it per line.

June 30, 2026·6 min read
Watercolor conversation between an author and an audio timeline
Product Updates

Direct Your Audiobook by Talking to It

Midsummerr's agentic director lets you edit a chapter in plain language — retime a pause, swap a voice, adjust a sound effect — and see exactly what changed.

July 15, 2026·3 min read

Midsummerr

Create premium audiobooks with cinematic quality in one click

[email protected]

Quick Links

HomeFeaturesServicesPricingAbout Us

Resources

BlogSupportRequest Demo

Legal

Terms of ServicePrivacy PolicyRefund Policy

© 2026 Midsummerr. All rights reserved.