TTS Narration Cleaner: Prepare Any Script for AI Voiceover
Paste a written script and get back narration a voice model can actually read: numbers spelled out, acronyms handled, citations and stage directions gone, sentences sized for breath.
Your script
0 / 15,000 characters≈ 0 sec of narration
Cleaned narration
Advanced options
Your original meaning is preserved. You control how much rewriting is allowed.
This tool improves spoken formatting. It does not fact-check your script, verify statistics, or confirm sources.
How to use it
- Paste the script, upload a .txt or .md file, or load the built-in example.
- Preserve My Wording changes pronunciation and formatting only. Improve Spoken Flow also shortens long sentences.
- Choose the narration type, then open Advanced options to control numbers, acronyms, paragraph size, and what gets removed.
- Press Prepare for TTS, read the change summary and the Words to Test list, then copy or download the narration.
Examples
Years and percentages
Acronyms and currency
Production notes
About this tool
Short version: paste your script, press Prepare for TTS, and copy the result into ElevenLabs, Google Cloud Text-to-Speech, Azure, PlayHT, Speechify, or any AI video tool. It's free, there's no sign-up, and the default mode never leaves your browser.
Scripts written to be read look fine on a page and fall apart out loud. Digits get guessed at, footnote markers turn into stray numbers, brackets become odd little pauses, and a 45-word sentence arrives in one breath. This tool rewrites those parts into the form you want spoken and leaves your meaning where you put it.
Preserve My Wording is deterministic and runs locally. Improve Spoken Flow adds a server-side AI pass that shortens and simplifies but is instructed never to add facts. Either way, the tool formats speech — it doesn't check whether what you wrote is true.
Frequently asked questions about TTS Narration Cleaner
Updated
- Does the tool change the meaning of my script?
- Not in the default mode. “Preserve My Wording” only makes pronunciation and formatting changes — numbers, acronyms, punctuation, markdown, and paragraph blocks. “Improve Spoken Flow” may reword sentences for easier listening, but it is instructed never to invent facts, add claims, or change your position. Always read the before-and-after comparison before rendering audio.
- Can I use it with ElevenLabs?
- Yes. Copy or download the cleaned narration and paste it into ElevenLabs Studio. This tool is independent and is not operated by, endorsed by, or affiliated with ElevenLabs in any official capacity.
- Does it work with other text-to-speech tools?
- Yes. The output is plain spoken text with no vendor-specific tags, so it works with Google Cloud Text-to-Speech, Azure, PlayHT, Speechify, podcast voice generators, audiobook tools, and AI video platforms.
- Why should numbers sometimes be written as words?
- Voice models guess at digits. “2026” can be read as “two thousand twenty-six” or “two zero two six”, and “$1.45 trillion” often loses the currency. Writing the intended spoken form removes the guess — which is why the tool converts years, currency, percentages, dates, and decimals into words.
- Does the tool remove citations and scene directions?
- By default it removes citations, footnote markers, markdown, headings, URLs, and bracketed production notes such as “[Show chart on screen.]”. Anything ambiguous — for example a long bracketed passage that looks like spoken dialogue — is kept and flagged for you to review instead of being deleted silently.
- Can I download the cleaned script?
- Yes. Copy it to the clipboard or download it as a .txt file. There is no account, sign-up, or watermark.
- Why does ElevenLabs read my numbers wrong?
- Because a digit has more than one spoken form and the model has to pick one. “2026” can come out as “two thousand twenty-six” or “two zero two six”, “1.45” can lose its decimal, and “$3M” often drops the currency or the magnitude. Writing the number as words is the only reliable fix, and it is the first thing this tool does.
- How do I make an AI voice spell out an acronym?
- Hyphenate it. “O-A-S-I” is read letter by letter in ElevenLabs, Google, Azure, and most other engines, while “NASA” left alone stays a word. The tool applies that rule automatically, lists every acronym it touched under Words to Test, and lets you override any of them.
- How long should a script be for text-to-speech?
- There is no hard rule, but shorter blocks cost less to fix. Splitting narration into one-to-three-sentence paragraphs means a single mispronounced word only costs you one block to regenerate, not the whole render. This tool rebuilds paragraph blocks for you based on the narration type you pick.
- Is my script stored?
- Faithful conversion runs entirely in your browser — nothing is uploaded. If you choose “Improve Spoken Flow”, the cleaned text is sent to our server, forwarded to the AI narration service, and held in a short-lived in-memory cache (about 15 minutes) so an identical repeat request is free. It is not written to a database, not attached to your identity, and never included in analytics.
- Does it check my facts?
- No. The tool improves spoken formatting only. It does not verify claims, statistics, dates, or sources — accuracy remains your responsibility.
How to prepare a script for AI narration
Why written scripts sound wrong out loud
Writing does a lot of work visually. Parentheses, semicolons, bullets, headings and citations all tell a reader's eye how to move. A voice model gets none of that. It clips brackets into strange pauses, reads footnote markers as loose numbers, and delivers a 45-word sentence without coming up for air. Stripping that structure out first usually improves the audio more than changing voices does.
Format numbers and dates the way you want them heard
Write the number as speech. “2026” is ambiguous; “twenty twenty-six” is not. Currency tends to lose its unit when a magnitude follows, so “$1.45 trillion” is safer as “one point four five trillion dollars”. Percentages, decimals, phone numbers and ordinals all have a preferred spoken form, and saying it explicitly removes the guess. Keep digits only where the digits are the point, like a version number the viewer also sees on screen.
Acronyms: spell out, or leave alone
Not every capitalised word should be spelled out. NASA is a word. IRS is three letters. SQL depends on who you ask. Hyphenating an initialism — “O-A-S-I” — gets letter-by-letter delivery from most engines, while leaving a lexicalised acronym intact keeps it as a word. Anything unusual belongs in a short test render before you spend credits on the full script.
Preserve wording, or let it be rewritten
Preserving your wording is the safe default: pronunciation and formatting get fixed, your sentences stay put. Rewriting is for scripts written for the page rather than the ear — long clauses, “however” and “furthermore”, layers of qualification. A good rewrite shortens and simplifies. If it hands back a fact you never wrote, throw it out.
Test before you spend voice credits
Render 30 to 60 seconds first, and pick the section with the most numbers and names in it. Listen for dates, dollar amounts, acronyms and proper nouns. Fix the text rather than the audio, then regenerate only the blocks that changed. Short paragraph blocks are what make that cheap — you never re-render a whole script because of one mispronounced word.
Which platforms this works with
The output is plain spoken text with no vendor tags, so it drops straight into ElevenLabs, Google Cloud Text-to-Speech, Azure Speech, Amazon Polly, PlayHT, Speechify, Murf, and the voiceover step in AI video tools. If your platform supports SSML, add those tags after cleaning — this tool deliberately leaves markup out so nothing gets read aloud by mistake.
Privacy
Everything runs locally in your browser. Your input is never uploaded, logged, or stored.