Choose the right model
Flash v2.5 targets ~75ms ultra-low latency excluding application and network latency. Multilingual v2 is the most stable for long-form content. Eleven v3 adds expressive, multi-speaker dialogue with audio tags. Match model to latency, language, and performance needs.
Stage directions get spoken
Text like “she said excitedly” influences delivery but is also spoken aloud. Trim those words in post if you only wanted the emotion, not the narration of it.
Instant vs Professional cloning
Instant Voice Cloning needs roughly 1–3 minutes of audio and returns in seconds. Professional Voice Cloning needs 30–180 minutes, a Creator plan or above, and hours of training—but yields the highest fidelity.
Voice Design when the library misses
Voice Design generates three previews from a 20–1000 character description plus optional preview text. Use it to explore voices; prefer an existing PVC in the library when one already fits.
Explicit language and dialect
Start Voice Design prompts with native language and regional variant. Being explicit reduces accent drift, especially for multilingual projects.
Avoid FX words in Voice Design
Terms like reverb, echo, phone, or tape can degrade quality unless you intentionally want degraded audio. Describe timbre and pacing instead.
Pauses and breaks
Use <break time="x.xs" /> up to 3 seconds on supported models. Eleven v3 does not support SSML break tags—use v3 prompting techniques instead. Too many breaks can destabilize output.
Pronunciation tooling
Eleven v3 accepts IPA wrapped in forward slashes. v2 Flash supports SSML phoneme tags. Studio and Dubbing accept pronunciation dictionaries (TXT or PLS) with alias or phoneme entries.
Long-form prosody
Split very long text and, on the API, pass previous_text and next_text so phrasing flows across chunks instead of resetting every segment.
Consistency and seeds
Outputs are nondeterministic. Use the seed parameter when you need repeatable takes, knowing minor variation may remain.
Free regenerations
You can regenerate the same text with identical voice settings up to two times at no extra cost—useful for occasional glitches, not for script changes.
Commercial rights
You own generated audio, but commercial usage requires a paid subscription. Verify your plan before monetizing outputs.
Voice Library on API
Community Voice Library voices are not available via API to free-tier users. Use voices in My Voices or upgrade accordingly.
Default voice sunset
Save production voice_id values you control (clones, designed voices, or library voices you added). Check ElevenLabs docs for voice availability on your plan.