Most people’s first ElevenLabs voiceover sounds like a satnav reading a tax return. Flat, faintly smug, and weirdly breathless at the end of every sentence. The fix rarely involves a better microphone or a more expensive tier. It comes down to about twenty small decisions you make before you ever press generate.
Here is the workflow I use on client work, broken into steps you can copy today.
Step 1: Choose the Model Before the Voice
ElevenLabs ships several models and they are not interchangeable. Picking the wrong one is the single biggest reason a good voice still sounds off.
- Eleven Multilingual v2 is the quality workhorse. Slower to render, best for narration, audiobooks and anything shipping as a finished file. It holds a consistent tone across 29 languages.
- Eleven Turbo v2.5 trades a little polish for speed. Fine for e-learning modules you will re-render a dozen times while the script changes.
- Eleven Flash v2.5 is built for latency, roughly 75ms, and exists for conversational agents rather than produced audio.
- English v1 is the older English-only model. Still handy when you need something cheap and quick.
Rule of thumb: if the audio will be heard more than once, use Multilingual v2 and walk away while it renders. If a user is waiting on a reply, use Flash.
Want the plan-by-plan breakdown of credits, character limits and commercial rights? There is a thorough ElevenLabs review that covers exactly that, and it is worth ten minutes before you hand over a card.
Step 2: Audition Voices Properly
The Voice Library holds thousands of community uploads, and Voice Design lets you describe a voice in plain text (“a warm mid-forties woman with a slight Irish accent, unhurried”). Both are useful. Neither replaces a proper audition.
Here is the method that actually filters them:
- Write one two-sentence sample that contains your hardest word and your longest sentence.
- Generate it with three candidate voices, same model, same settings.
- Listen at normal speed on laptop speakers, not headphones. Most of your audience hears it worse than that.
One voice will usually nail the hard word and mangle nothing else. That is your pick, and it is worth saving as a preset.
Step 3: Set Stability and Similarity, Then Leave Them Alone
Two sliders do most of the work.
Stability controls how much the delivery varies. Low values, roughly 20 to 35, give you range: a voice that leans on words, drops to a near-whisper, sounds like it means what it says. High values, 70 and up, give you a calm, repeatable read. High stability is also where the dreaded monotone lives.
Similarity controls how closely the output tracks the source voice. Push it too far and you inherit artifacts from the original recording. Around 75 is a solid default. Move to 85 only when a cloned voice starts drifting between renders.
Starting points that hold up in practice:
- Documentary or YouTube narration: stability 50, similarity 75
- Character dialogue or emotional ad copy: stability 28, similarity 70
- Phone system or in-app assistant: stability 75, similarity 80
- Cloned brand voice that must sound identical every time: stability 65, similarity 85
Set them once. Re-tweaking after every sentence is how a twenty-minute job becomes a Tuesday.
Step 4: Rewrite the Script for a Voice Model
Text written for eyes fails in a mouth. A handful of edits fix most of the clunk:
- Spell out numbers and symbols. “£4.2m” becomes “four point two million pounds”.
- Expand abbreviations on first use. “approx.” becomes “approximately”.
- Kill ellipses and replace them with commas or full stops. Trailing dots send delivery in odd directions.
- Chop forty-word sentences into two. Long clauses are where pacing falls apart.
- Add commas where you want a breath, not where grammar insists on one.
You can also insert explicit pauses. ElevenLabs supports break tags, so a line like We shipped it in March. <break time=”0.6s” /> Then everything broke. lands the joke instead of racing past it.
Step 5: Handle Names With Phonemes
Proper nouns are where a good read goes sideways. Two approaches, both fast.
The low-tech one: respell the name the way it sounds. “Nguyen” becomes “New-win”, “Siobhan” becomes “Shiv-awn”. Generate a ten-second test, check it, move on.
The precise one: wrap the word in a phoneme tag using CMU Arpabet. That gives you exact control, which matters when you are producing forty episodes and cannot re-check every guest name by ear.
Watch heteronyms too. “Lead” the metal and “lead” the verb are the same string of letters. If context is ambiguous, swap the word out entirely.
A Sixty-Second Example, Start to Finish
Raw script: “Our new app saves teams 4.5 hrs/wk — approx. 20% of a working week — by automating the boring bits.”
Rewritten for the model: “Our new app saves teams four and a half hours a week. That is roughly twenty percent of a working week. All of it comes back to you, because the boring parts now happen on their own.”
Same message, more words, far better audio. The em dash became a full stop, the shorthand became speech, and the jargon got unpacked. Run that through Multilingual v2 at stability 50 and it sounds like a person who has read the sentence before opening their mouth.
Going Further: Batch Work, Avatars and Self-Hosting
Once a single clip sounds right, two things usually happen.
The first is scale. Projects handles long-form chapters with consistent settings, and the API lets you queue hundreds of clips from a spreadsheet. For anything scripted on camera, matching that audio to a face is a separate problem, and this breakdown of what AI avatars can and can’t do for your video will save you from a stack of re-shoots.
The second is privacy. If your scripts cannot leave your own machine, ElevenLabs is the wrong tool, and a self-hosted route such as cloning a voice with Coqui AI is worth the setup time.
If budget is the barrier before you get that far, free tiers are more generous than most people assume, though plenty of them are not worth the afternoon. This rundown of free AI tools worth using in 2025 separates the ones that actually work from the ones that stall at the signup screen.
The Five Failures You Will Hit, and the Fix for Each
- It rushes. Raise stability by 10 and add commas at clause boundaries.
- It sounds robotic. Lower stability to 30 and re-render on Multilingual v2 instead of Turbo.
- One word keeps coming out wrong. Respell it or use a phoneme tag. Re-rendering alone will not fix it.
- It drifts partway through a clone. Raise similarity to 85 and split the script into smaller chunks.
- It sounds different every time. You are changing settings between renders. Lock them and edit only the text.
Save a preset per project once you find a combination you like, and the whole workflow collapses to paste, generate, ship.

