How to Create Realistic AI Voices With ElevenLabs: Step-by-Step Tutorial (2026)
ElevenLabs tutorial for 2026: choose v4 or v3, find a voice, use model-specific settings and audio tags, prepare your script, clone your voice and export.
Published . Updated . 15 min read. Based on official sources, checked on the update date; every price links to its source. Our method.
Contents
- What You Need Before You Start
- Step 1: Pick the Right Model for the Job
- Step 2: Find a Voice in the Voice Library
- Step 3 (Optional): Create a Voice With Voice Design
- Step 4: Tune Stability, Similarity, Style and Speed
- Step 5: Direct Eleven v3 With Audio Tags
- Step 6: Prep Your Script So the Voice Reads It Right
- Step 7: Clone Your Voice With Instant Voice Cloning
- Step 8: Generate, Review and Export
- Commercial Rights: What Each Plan Allows
- Common Problems and How to Fix Them
- FAQ
By the end of this tutorial, you’ll have a finished voiceover generated from your own script in ElevenLabs, using a voice you picked or designed for the project, directed with the appropriate model, and downloaded. It takes eight steps. A free account is enough to practice (about 10 minutes of speech a month, non-commercial only). To publish a monetized video or clone your voice, you need at least the Starter plan at $6 a month plus tax (pricing, checked Sept. 25, 2026).
Model selection was updated for Eleven v4 on September 28, 2026; v3-specific settings and examples are labeled below. For full pricing, the rest of the platform (dubbing, music, voice agents) and a comparison with Murf, see our ElevenLabs guide with pricing and free plan details.
What You Need Before You Start
Create an account on ElevenLabs (affiliate link: we earn a commission if you subscribe, at no extra cost to you). Every feature draws on the same credit balance. For text to speech, 1 character costs 1 credit, so a minute of speech runs about 1,000 credits.
What matters for this tutorial, plan by plan (pricing page, checked Sept. 25, 2026):
| Plan | Monthly price (plus tax) | Speech per month | Commercial use | Instant Voice Cloning | Audio export |
|---|---|---|---|---|---|
| Free | $0 | about 10 min | No, and attribution required | No | MP3 128 kbps |
| Starter | $6 | about 30 min | Yes | Yes | MP3 128 kbps |
| Creator | $22 ($11 first month) | about 121 min | Yes | Yes, plus Professional Voice Cloning | MP3 128 and 192 kbps |
| Pro | $99 | about 600 min | Yes | Yes | MP3 128 and 192 kbps, WAV/PCM 44.1 kHz |
Limits depend on the model and interface: the model documentation lists 10,000 characters for v4 and 5,000 for v3. Also check the counter in your Text to Speech workspace; use Studio to manage longer projects in sections (models, checked Sept. 28, 2026).
Step 1: Pick the Right Model for the Job
You choose the model in the settings on the Text to Speech page. According to the Text to Speech guide, it’s the second biggest factor in quality after the voice itself.
| What you’re making | Model | Why |
|---|---|---|
| New narration project or professional voice clone | Eleven v4 | Current generation, 90+ languages, up to 10,000 characters |
| Real-time conversation | Eleven v4 Turbo | v4 variant for conversational use |
| Existing projects or the tag examples below | Eleven v3 | Expressive, directed with audio tags, 70+ languages, 5,000 characters per generation |
| Audiobooks, long scripts, very consistent delivery | Multilingual v2 | “Most stable on long-form generations,” per ElevenLabs; 29 languages; supports <break> pauses |
| Real-time apps, voice assistants, high API volume | Flash v2.5 | Around 75 ms latency, 32 languages, half the credits per character through the API |
Eleven v3 isn’t built for real time: its latency is higher and its output varies more from one take to the next. The older Turbo v2/v2.5 models were replaced by Flash; v4 Turbo is a separate new model (models documentation, checked Sept. 28, 2026).
For a new project, try v4 on a few sentences and compare it with your previous model before switching a whole narration. v4 and v4 Turbo launched on September 28, 2026 (official announcement). The settings and tag examples below identify the instructions specific to v3.
Step 2: Find a Voice in the Voice Library
The voice matters, but the accent also depends on the model. With v4, ElevenLabs says speech adapts to a native accent in the target language, even when the source voice speaks another language (v4 announcement, checked Sept. 28, 2026). Listen to a short sample in your target language before generating a full script.
To find a voice with a specific accent or style (Voice Library guide):
- In the sidebar, open Voices, then Explore: this is the library of voices shared by the community.
- Filter by Language (English, for example), then by Accent if you want a specific one, such as American.
- Narrow it down by category (Narration, Advertisement, Educational…), gender and age.
- Listen to the previews, then click Use voice: Text to Speech opens with that voice selected. The + button saves it to My Voices for later.
Two things to watch for. Some voices are only available on paid plans (“This voice is not available for free users”). Others carry a credit multiplier, shown as a tag: with a 2x multiplier, every generation costs twice as many credits (help center, checked Sept. 25, 2026).
Step 3 (Optional): Create a Voice With Voice Design
If nothing in the library fits, Voice Design creates a voice from a written description, even on the free plan. The path in the docs: Voices, My Voices, add a new voice, then Voice Design (Voice Design guide, checked Sept. 25, 2026).
ElevenLabs recommends structuring the description like this: native language and variant, gender and age, audio quality, persona, emotion, then one or two sentences on timbre and pacing. An example to adapt:
Native English, neutral American accent. Male, 40–50. Studio quality.
Persona: documentary narrator. Emotion: calm, warm, confident.
Deep, warm timbre with a measured pace, crisp articulation and short pauses between sentences.
And a preview text that matches the voice (longer previews in the right tone give more stable results):
Twenty thousand years ago, these cliffs already towered over the valley. Even today, the wind tells the same story… to anyone who takes the time to listen.
Each click on Generate voice gives you three options and only charges for the characters in the preview text. Saving a voice uses a custom voice slot: 3 on the free plan, 10 on Starter, 30 on Creator (pricing, checked Sept. 25, 2026). The docs advise against effect words (“reverb,” “phone,” “tape”) and against writing “accent” when you mean intonation, which can trigger an unwanted dialect shift. Voices made with Voice Design v3 work with Eleven v3 and its audio tags.
Step 4: Tune Stability, Similarity, Style and Speed
Available controls depend on the model. Start with the interface defaults for v4; the table below covers classic v2/Flash controls, followed by v3-specific settings. Starting values and effects from the Text to Speech guide (checked Sept. 25, 2026):
| Setting | What it does | Suggested starting point |
|---|---|---|
| Stability | Lower: more expressive but unpredictable, sometimes too fast. Higher: consistent, up to monotone | Around 50 |
| Similarity | How closely the output sticks to the original voice. Too high with a noisy source reproduces the noise and artifacts | Around 75 |
| Style exaggeration | Amplifies the original speaker’s style, at the cost of stability | 0 |
| Speaker boost | Pushes a bit closer to the original speaker; subtle effect | Your call |
| Speed | From 0.7 (slower) to 1.2 (faster), 1.0 by default | 1.0; extreme values hurt quality |
Eleven v3 works differently. Stability becomes a choice between three modes (best practices guide):
- Creative: the most expressive, but prone to hallucinations;
- Natural: closest to the original voice, balanced;
- Robust: very stable, similar to v2, but less responsive to audio tags.
The docs state that Similarity, Speaker boost and Speed aren’t available for Eleven v3, so you control pacing with audio tags and punctuation instead. For a tag-directed voiceover, start on Natural and move to Creative if the read still sounds flat.
No two generations are identical, even with the same settings. Generate two or three takes and keep the best one.
Step 5: Direct Eleven v3 With Audio Tags
Eleven v3 takes audio tags in square brackets, placed right in your script. They set emotion ([excited], [sad], [curious]), delivery ([whispers], [shouts]), reactions ([laughs], [sighs], [clears throat]) and pacing ([pause], [slowly], [rushed]) (best practices guide and the Audio Tags 101 post, checked Sept. 25, 2026).
Ready-to-paste examples for Text to Speech with Eleven v3 selected.
YouTube intro:
[excited] Welcome back to the channel! Today we're making a full voiceover in under ten minutes. [pause] Ready? Let's go.
Audiobook narration:
[softly] The house had been silent for years… [pause] And yet, that night, someone had turned on a light in the attic. [whispers] She wasn't alone.
Ad with a change in tone:
[sighs] Another electric bill that keeps going up? [pause] [excited] With our calculator, you'll know in two minutes how much you could save. [slowly] Two minutes. That's it.
Customer service, professional then reassuring:
[professional] Hi, and thanks for calling. [sympathetic] I understand, a late delivery is really frustrating. [reassuring] Let me check where your package is right now.
Three rules from the docs:
- The voice has to be able to play the tag: a calm, meditative voice won’t shout convincingly with
[shouts], and a whispering voice won’t suddenly yell. - No
<break>with Eleven v3: for pauses, use[pause], ellipses or a line break. The<break time="1.5s" />tag (up to 3 seconds) still works with Multilingual v2 and Flash v2.5. - Capital letters add emphasis in Eleven v3: “It was a VERY long day.”
The Enhance button in the interface adds audio tags to your text automatically. It’s a good starting point, but review what it adds.
Step 6: Prep Your Script So the Voice Reads It Right
The voice reads exactly what you write, typos included: it doesn’t correct anything. Your script matters as much as the settings (troubleshooting guide, checked Sept. 25, 2026).
Numbers, dates, symbols. Write them out the way they should be spoken. The website normalizes numbers automatically, but ElevenLabs still recommends spelling them out, and smaller models can misread them: its docs note that Flash v2.5 reads “$1,000,000” as “one thousand thousand dollars,” where Multilingual v2 gets it right (best practices guide).
Abbreviations and acronyms. Expand abbreviations (“Dr.” becomes “Doctor,” “Ave.” becomes “Avenue”) and write acronyms the way you want them said.
Punctuation. Ellipses add a pause plus some hesitation; a dash gives a more neutral pause; short sentences make the pacing sound more natural.
A before-and-after to paste and compare:
Raw script:
On 09/25/2026 at 2:30 PM, Dr. Smith paid $6 for 30 min of audio, 20% less than last year. Call 555-123-4567 or visit example.com/help.
Prepped script:
On September twenty-fifth, twenty twenty-six, at two thirty PM, Doctor Smith paid six dollars for thirty minutes of audio. That's twenty percent less than last year. Call five five five, one two three, four five six seven, or visit example dot com slash help.
Names and brand terms. With Eleven v3, you can force a pronunciation with International Phonetic Alphabet (IPA) symbols wrapped in forward slashes, as in the docs’ examples. ElevenLabs reports 80 to 90% pronunciation consistency, so check every take.
Our next stop is "/ˌsænfrənˈsɪskoʊ/", on the West Coast.
Otherwise, respell the word the way it sounds. In Studio, a pronunciation dictionary (TXT or PLS file) applies your rules to a whole project (pronunciation best practices, checked Sept. 25, 2026).
Length. Split a script at scene or topic boundaries to make corrections easier. The older recommendation to stay below 800–900 characters is not a v4 limit: the model documentation lists up to 10,000 characters per request (checked Sept. 28, 2026). For an audiobook, use Studio and listen to transitions between sections.
Step 7: Clone Your Voice With Instant Voice Cloning
Instant Voice Cloning is included from the Starter plan up. Steps from the official guide (checked Sept. 25, 2026):
- In Voices, click the + icon, then choose Instant Voice Clone.
- Upload or record your audio, following the on-screen instructions.
- Name the voice, confirm you have the right and consent to clone it, then click Save voice.
- In My Voices, click Use voice to use it in Text to Speech.
What makes a good clone:
- 1 to 2 minutes of clean audio, no more than 3: beyond that, the gain is small and the clone can become less stable;
- one speaker only, with no music, background noise or room echo;
- consistent tone and mic distance, no whispering or shouting; average level between -23 and -18 dB RMS, peaks below -3 dB;
- MP3 at 192 kbps or higher: WAV doesn’t improve the clone;
- the delivery you want back: the clone copies your pace and energy. Record in the language and accent you want it to speak.
A clone can’t be exported: it stays in your account, and the API is the only way to use it elsewhere. If your accent doesn’t come through, change the samples or move up to Professional Voice Cloning (Creator plan and up, covered in our ElevenLabs guide). Eleven v4 supports both instant and professional clones. For Eleven v3 audio tags, ElevenLabs recommends an Instant Voice Clone or a Voice Design voice, since Professional Voice Clones aren’t fully optimized for that model yet.
Consent. The usage policy prohibits cloning another person’s voice without their consent or a legal right. For a coworker, a voice actor or a brand voice, get a signed written agreement before you record.
Step 8: Generate, Review and Export
- Click Generate speech. Every generation is charged, whether or not you download it.
- Listen. If it’s not right, you get two free regenerations as long as you keep the same text, voice and model (you can move the sliders), within two hours and without refreshing the page. The button then reads “Regenerate speech” (help center, checked Sept. 25, 2026).
- Download the file with the download button at the bottom right after generating, or later from the history tab in the right-hand panel.
From your history, the download menu offers MP3 or WAV, and Advanced unlocks other formats (MP3 at 192 or 256 kbps, M4A, FLAC) (help center, checked Sept. 25, 2026). The quality you get depends on your plan: per the comparison table on the pricing page, MP3 at 128 kbps on Free and Starter, 192 kbps from Creator up, and WAV/PCM 44.1 kHz from Pro up (checked Sept. 25, 2026).
To put your voiceover to work, see our tutorials on making an educational video with Synthesia and ChatGPT and composing music with AI.
Commercial Rights: What Each Plan Allows
ElevenLabs’ publishing rules (help center, checked Sept. 25, 2026):
- Free plan: no commercial use at all. Anything you publish must include “elevenlabs.io” or “11.ai” in its title.
- Paid plans: commercial license included, indefinitely, for content generated during your subscription, as long as you hold the rights to the text and voice you used.
- No retroactive license: audio generated on the free plan (before or after a subscription) can never be used commercially. Regenerate it once you’ve subscribed.
- Beta services excluded: their output can’t be used commercially or in production.
A monetized YouTube video, an ad or client work counts as commercial use, so plan on Starter at a minimum.
Common Problems and How to Fix Them
The Voice Sounds Metallic or Full of Artifacts
- If a long passage develops artifacts, isolate and regenerate the affected sentence; 800 characters is not a universal limit.
- Style exaggeration is above 0: ElevenLabs recommends leaving it at 0, since it can add stray sounds and uneven pacing.
- With a cloned voice, Similarity set too high reproduces flaws in the source recording: lower it, or redo the clone with cleaner audio.
- Too many
<break>tags in one text can speed up the voice and add artifacts. - Muffled or distorted audio can happen at random: regenerate that section (for free, if the conditions in Step 8 apply).
The Accent Is Wrong or Drifts
- The voice wasn’t recorded with the accent you want: filter the library by language and accent, or clone a voice that has it.
- For a long or multilingual passage, compare a shorter excerpt using the same model to isolate accent drift.
- In Voice Design, start the description with the language and variant (“Native English, neutral American accent”).
- If a voice picks up another accent in multilingual content, write numbers and acronyms out in the target language.
Odd Pauses, Cut-Offs or Rushed Delivery
- Stability set too low makes the voice unpredictable and fast: move it back toward 50, or switch to Natural on Eleven v3.
- A voice cloned from short, choppy clips inherits that rhythm: ElevenLabs recommends longer, continuous samples.
- Ellipses add hesitation: for a neutral pause, use a dash,
[pause]on Eleven v3, or<break>on Multilingual v2. - In Studio, a glitch or sharp breath between paragraphs usually comes from the paragraph before it: regenerate that one.
- If the interface counter rejects your script, split it or use Studio; the limit depends on your model and plan.
Credits Disappear Too Fast
- Every click on Generate is charged, download or not. Dial in the voice and settings on one or two sentences before running the full script.
- Use your two free regenerations: same text, same voice, same model, sliders can change.
- Check for a credit multiplier tag before adopting a Voice Library voice.
- Voice Design only charges for the preview text characters, once for all three options.
- On a paid plan, unused credits roll over for up to two months; the free plan doesn’t roll anything over (pricing, checked Sept. 25, 2026).
FAQ
Is ElevenLabs Free for Voiceovers?
Yes, for practice: the free plan gives you about 10 minutes of speech a month, with Eleven v3 and Voice Design, but no cloning or commercial use, and anything you publish needs “elevenlabs.io” or “11.ai” in its title. Starter costs $6 a month plus tax (checked Sept. 25, 2026 on the pricing page).
How Do I Make ElevenLabs Speak Another Language?
Just write your script in that language: the model detects the language from the text, and there’s no language selector on the website. To avoid an English accent, pick a voice tagged with that language in the Voice Library, or clone a native speaker (with their consent).
Which Model Should I Use for a Voiceover?
Start a new project with Eleven v4; select v3 to follow the specific audio-tag examples in this tutorial. v4 Turbo targets real-time conversations, while Multilingual v2 remains an option for existing long-form projects.
Can I Clone Someone Else’s Voice?
Only with their consent or a legal right: ElevenLabs’ usage policy requires it, and the app asks you to confirm it before saving the clone. Get it in writing.
Is There an Alternative for Professional Voiceover?
Murf Studio targets e-learning and presentation voiceover. Canva is included from Creator; PowerPoint and Google Slides require Business. Custom voice cloning is an Enterprise add-on, distinct from translating a recording in Murf Dub (pricing, checked Sept. 30, 2026). The side-by-side numbers are in our ElevenLabs guide. To publish your voiceover as a show, see our tutorial on automating a podcast.