Audie
HomePricing

Chapter Detection: Splitting Your Manuscript for Audiobook Production

Audie
By Audie
August 11, 2026

If you have poured months into a manuscript, produced your audiobook, and you are now hovering over the ACX upload button wondering whether your files will actually pass, take a breath. This is the part where a lot of authors get nervous, and honestly, it is the part with the clearest rules of the whole journey. The ACX audio submission requirements come down to a short, specific list of technical targets: your files need to sit between roughly -23 dB and -18 dB RMS, peak no higher than -3 dB, keep a noise floor quieter than -60 dBFS, be 192 kbps CBR MP3 at 44.1 kHz, and open with the required opening and closing credits. Hit those numbers and your book passes review. Miss one and it bounces back with a note. That is really the whole game.

I want to walk you through each requirement plainly, because the terms sound more intimidating than they are, and because AI-narrated production changes which of these you actually have to worry about. When you generate your audio with clean neural voices instead of recording in a spare bedroom, half the classic ACX rejections simply never happen. Let me show you what to check, what to ignore, and where the real gotchas hide.

The core ACX audio submission requirements, in plain numbers

ACX (Audiobook Creation Exchange, Amazon's audiobook production and distribution platform for Audible) publishes a fixed set of technical specs every file must meet. Here is the full checklist, translated out of engineer-speak:

  • RMS (loudness) between -23 dB and -18 dB. RMS is the average, perceived loudness of your file, not the loudest single moment. Too quiet and listeners crank their volume; too loud and it clips or fatigues the ear. ACX wants your overall level to live in that window.
  • Peak level at -3 dB or lower. The peak is the single loudest instant in the file. Leaving 3 dB of headroom below the digital ceiling (0 dB) prevents clipping, that ugly crackle when audio slams into the maximum. Think of it as a safety margin.
  • Noise floor quieter than -60 dBFS. The noise floor is the level of the "silence" between words: the hiss, hum, or room tone underneath everything. Human recordings struggle here. Clean AI-generated audio is effectively silent in the gaps, so this one is usually a non-issue.
  • 192 kbps constant bit rate (CBR) MP3, 44.1 kHz sample rate. This is the file format ACX accepts. Constant bit rate, not variable. 44.1 kHz is CD-quality sampling.
  • Each chapter as its own file, under 120 minutes each, with a matching, sensible file name.
  • Opening and closing credits, plus a retail sample, which I will cover in its own section because it is where people slip up.
  • 0.5 to 1 second of room tone at the start, and 1 to 5 seconds of silence at the end of each file. Not dead-digital silence necessarily, but a clean, quiet tail.

That is the entire technical bar. Notice how much of it is about consistency and cleanliness rather than fancy equipment. If you want the bigger-picture view of whether AI narration is even allowed on the platform before you sweat the specs, read Audie's guide to whether ACX allows AI-narrated audiobooks under the current 2026 policy first: the answer has shifted, and it matters.

Audie

Check the policy before you check the specs. There is no point mastering perfect files for a platform that will not accept your production method. Read the policy piece first.

RMS, peak, and noise floor: what they mean and how to hit them

These three are the numbers that trip up first-time producers, so let me sit with them for a minute. They all measure different things, and understanding the difference is what lets you fix a rejection instead of guessing.

RMS is your average loudness. Imagine measuring the energy of your whole file and flattening it into one number. ACX wants that number between -23 and -18 dB. If ACX tells you the file is "too quiet," your RMS is below -23. The fix is a gentle overall gain boost, or normalization to a target like -20 dB RMS, which lands you comfortably in the middle of the window. If it comes back "too loud," you pull that same gain down. Most audio editors, including free ones like Audacity, have a Loudness Normalization tool that does this in one pass.

Peak is your single loudest moment. You want it at -3 dB or below. Here is the subtle part: RMS and peak are related but not the same. You can have a correct average loudness while one dramatic shout or hard consonant spikes past -3 dB. The clean fix is a limiter set to a -3 dB ceiling, which catches only the peaks that poke through and leaves everything else alone. Do the limiting first, then normalize RMS, in that order.

Noise floor is your silence. ACX requires it below -60 dBFS, meaning the quiet gaps between words must be genuinely quiet. This is the requirement that historically forced authors into closets full of blankets, USB microphones, and hours of noise reduction. A furnace kicking on, a laptop fan, traffic outside: all of it lives in the noise floor and all of it fails review. And this is exactly where AI narration quietly wins. When you generate audiobook audio from clean neural voices, the "silence" is actual digital silence. There is no room, no fan, no hiss to reduce. Your noise floor is effectively perfect before you touch a single setting.

Audie

The noise floor is where I have watched the most authors lose weeks. It is such a relief that clean AI audio simply does not have that problem. One less closet full of blankets.

Credits, samples, and file structure: the requirements people forget

Here is a mistake I see all the time, and it has nothing to do with loudness numbers. Authors master a beautiful set of files and then get bounced because they forgot the credits or structured the files wrong. These are easy to fix once you know to look.

Opening credits go at the very start of your first file and must state the title, the author's name, and the narrator's name. For AI production, ACX's current guidance is to disclose the AI narration in that credit, which is straightforward to add. Closing credits come at the end of the final file and typically restate the title and add "The End" or a production line.

The retail sample is a 1 to 5 minute clip ACX uses as the preview on your Audible product page. It has to be representative audio pulled from the book itself, not the opening credits, and it must meet all the same technical specs. Pick a passage that shows off your narration at its best, ideally one with a bit of character voice or momentum so a browsing listener wants to hear more. If you are producing with distinct voices per speaker, a scene with dialogue makes a strong sample: Audie's guide to giving every character a different voice in your audiobook walks through how that works.

File structure matters too. Every chapter is a separate MP3, none longer than 120 minutes, named clearly and in order (something like "01_Prologue," "02_Chapter-One," and so on). ACX also requires a separate opening credits file and closing credits file. If you generate your book with a tool that already splits by chapter and hands you cleanly named, chaptered MP3s, this whole section becomes drag-and-drop. If you are wrestling one giant three-hour file, splitting it correctly is its own chore, which is why Audie's walkthrough on chapter detection and splitting your manuscript for production is worth a look before you upload.

One more quiet detail: ACX wants 0.5 to 1 second of room tone at the head of each file and 1 to 5 seconds of clean silence at the tail. For a human recording, "room tone" means a recorded snippet of your actual room's quiet ambience so the start does not cut in jarringly. For clean AI audio, a short lead-in and a proper silent tail accomplish the same thing without any recorded hiss to match.

Why AI-narrated files pass review more easily

I want to be candid here, because I do not think enough people connect these dots. A huge share of ACX rejections are technical, not artistic: inconsistent loudness between chapters, a noise floor that creeps above -60 dBFS in one file, a peak that clips during an emotional scene, mouth clicks and breaths that were never cleaned up. Every one of those is a home-recording problem.

When you produce your audiobook with Audie, you are generating from Azure HD neural voices and ElevenLabs voices that output clean, consistent digital audio. The loudness is uniform from chapter one to chapter forty. The noise floor is silence. There are no plosives to de-ess, no fan to filter, no accidental clipping from a narrator leaning into the mic. You start much closer to spec than a home recording ever could, which means the mastering step is a light touch rather than a rescue mission. Audie produces downloadable, chaptered MP3s with automatic speaker detection and per-character voice assignment, so the file structure ACX wants is close to what you already have in hand. If you want the full walkthrough of that path, start with Audie's flagship guide on how to turn your book into an audiobook with AI.

That said, I will not pretend the files come out pre-stamped with an ACX seal. You still want to run a final loudness normalization to land inside that -23 to -18 dB RMS window and confirm your peak sits at -3 dB, because target loudness is a mastering choice, not something any generator guarantees. But you are doing a five-minute polish instead of a five-hour salvage. If you are weighing the full cost and effort of this path against hiring a narrator, Audie's comparison of AI narration versus hiring a human narrator on cost, time, and quality lays the tradeoffs out honestly.

Audie

Clean audio in, easy approval out. Run one loudness pass, add your credits, and you are genuinely ready to submit. It is not the ordeal it used to be.

Your pre-submission checklist

Before you upload anything to ACX, run down this list. If every box is checked, your files should sail through review:

  • RMS loudness is between -23 dB and -18 dB (aim for around -20 dB).
  • Peak level is at -3 dB or lower, with no clipping.
  • Noise floor is quieter than -60 dBFS in every file.
  • Every file is 192 kbps CBR MP3 at 44.1 kHz.
  • Each chapter is a separate file, under 120 minutes, clearly named in order.
  • Opening credits file includes title, author, narrator, and AI disclosure.
  • Closing credits file is present.
  • A representative 1 to 5 minute retail sample is prepared and meets spec.
  • Each file has 0.5 to 1 second of head room tone and 1 to 5 seconds of clean tail silence.
  • Loudness and levels are consistent across all chapters.

Print it, keep it beside you, and you will not be guessing when you hit submit. And if you would rather skip most of this by starting with clean, chaptered, consistent audio in the first place, that is exactly what Audie is built for.

Audie

Come make yours with me. You will get clean, chaptered MP3s that start most of the way to ACX-ready. See the pricing and give it a try.

Frequently asked questions

What RMS and peak levels does ACX require?

ACX requires your files to measure between -23 dB and -18 dB RMS (average perceived loudness) with a peak level no higher than -3 dB. Aiming for roughly -20 dB RMS puts you in the middle of the accepted window, and a limiter set to a -3 dB ceiling keeps your peaks safely below clipping. Run limiting first, then loudness normalization.

What is the noise floor requirement and why do AI voices pass it easily?

ACX requires a noise floor quieter than -60 dBFS, meaning the silence between words must be genuinely quiet with no hiss, hum, or room tone. Home recordings struggle here because fans, furnaces, and traffic all live in that quiet. AI-generated audio from clean neural voices has effectively digital silence in the gaps, so it passes the noise floor requirement without any noise reduction at all.

What file format does ACX accept for audiobook submissions?

ACX accepts 192 kbps constant bit rate (CBR) MP3 files at a 44.1 kHz sample rate. Each chapter must be a separate file no longer than 120 minutes, clearly named in sequence, alongside a separate opening credits file, closing credits file, and a retail sample. If your production tool delivers chaptered MP3s, most of this structure is already handled.

Do AI-narrated audiobook files still need mastering before ACX submission?

Yes, but only a light touch. Clean AI audio starts close to spec with a silent noise floor and consistent loudness, so you avoid the heavy repair work home recordings need. You still want a final loudness normalization to land inside the -23 to -18 dB RMS window and a check that your peak sits at -3 dB, since target loudness is a mastering choice rather than something any generator guarantees.

Audie is your guide to making audiobooks with AI at audie.ai. She helps authors turn their manuscripts into professional multi-voice audiobooks: no studio, no fuss.

Audie

Audie

Audie is your guide to making audiobooks with AI at audie.ai. She helps authors turn a manuscript into a professional multi-voice audiobook - no studio, no fuss.

Create Your Own AI Audiobook

Transform your manuscript into a professional audiobook in minutes with cutting-edge AI voice technology.

Audie
Made withfor authors
© 2026 Audie AI