Skip to content
ElevenLabs logo

ElevenLabs

AI voice generator and voice agents platform for lifelike speech, cloning, and conversational agents.

AudioAgents

How it works / How to use

AI voice generator and voice agents platform for lifelike speech, cloning, and conversational agents. Pick or create a voice, paste your script, and generate audio. Models read emotional cues from the text itself; Stability and Similarity settings control consistency. Instant cloning works from short samples; Professional Voice Cloning needs extended training audio on a Creator plan or above.

  1. Open ElevenLabs and choose Text to Speech, Voice Design, or Voice Cloning depending on the job.
  2. Select a voice from the Voice Library, a clone you own, or design a new voice from a text description.
  3. Paste or type your script. Add emotional cues in the text when delivery matters.
  4. Adjust voice settings (Stability, Similarity boost, and model choice), generate, then refine one variable at a time.

Text-to-speech narration

Paste the final script, choose a model (Flash v2.5 for low latency, Multilingual v2 for long-form quality, Eleven v3 for expressive delivery), then generate.

What to provide

  • The exact script to speak, including punctuation and emphasis
  • Target language and accent (match voice to region when possible)
  • Use case: ad, audiobook chunk, IVR line, or product demo

Details that improve the result

  • Models interpret emotion from text—phrases like “she said excitedly” or exclamation marks influence delivery
  • Descriptive stage directions in the text will be spoken unless you trim them from the audio afterward
  • For long text, split into segments and use previous_text / next_text (API) to keep prosody natural across chunks

Example prompt

Voice: calm female narrator, neutral American accent.
Script: Welcome back. Today we walk through the setup in three steps—nothing fancy, just what you need to ship.
Model: Multilingual v2 for a polished explainer. Keep pacing measured.

If the first output is not good

  • Regenerate the same text once if quality glitches (up to 2 free regenerations with identical settings)
  • Lower Stability slightly for more emotional range; raise it for a steadier read
  • Shorten one paragraph and ask for tighter pacing without rewriting the whole script

Common mistakes

  • Expecting stage directions to stay silent—they are usually spoken
  • Using one giant block of text instead of splitting long-form content
  • Choosing a voice whose accent does not match the script language

Instant voice cloning

Go to Voices → add a new voice → Instant Voice Cloning, upload samples, then test with a short TTS generation.

What to provide

  • 1–3 minutes of clean audio of the voice you have rights to clone
  • Consistent recording quality (MP3 at 192 kbps or higher recommended for PVC; IVC works from shorter samples)
  • A test script that reflects how the clone will be used

Details that improve the result

  • IVC produces results in seconds from about 1–3 minutes of audio
  • Works well for most voices; uncommon accents or very unique voices may need Professional Voice Cloning instead
  • Instant Voice Clones cannot be shared publicly in the Voice Library

Example prompt

Upload two minutes of dry studio narration. Test line: “This is a quick sample to check clarity and pacing before we record the full module.” Compare against the source for accent drift.

If the first output is not good

  • If the clone misses your accent, switch to Professional Voice Cloning with more training audio
  • Adjust Similarity boost to stay closer to the source voice
  • Re-upload cleaner samples without background music or room echo

Common mistakes

  • Training on noisy, compressed, or multi-speaker audio
  • Cloning a voice you do not own or have permission to use
  • Expecting IVC to match PVC fidelity for niche accents

Professional voice cloning

Create a Professional Voice Clone, upload or record extended training audio, wait for fine-tuning (often 3–6 hours), then validate across several test scripts.

What to provide

  • 30–180 minutes of high-quality training audio you own
  • Creator plan or above (required for PVC)
  • Scripts aligned to your intended use (narration, conversational, or advertising samples help)

Details that improve the result

  • PVC trains dedicated models for fidelity that is virtually indistinguishable from the source
  • Recommended format: MP3 at 192 kbps or higher
  • PVCs auto-train on Flash v2.5, Turbo v2.5, and Multilingual v2; English PVCs also include Flash v2 and Turbo v2

Example prompt

Upload 45 minutes of podcast narration recorded in the same room. After training completes, test: (1) a calm explainer, (2) an upbeat promo, (3) a sentence with numbers and product names.

If the first output is not good

  • Add more varied training audio if emotional range is flat
  • Use pronunciation dictionaries for brand names and acronyms
  • Compare PVC output against IVC on the same script before committing

Common mistakes

  • Submitting inconsistent mic setups across training files
  • Assuming PVC is instant—it requires training time
  • Skipping voice-captcha verification requirements

Voice design from a prompt

Use Voices → My Voices → Add a new voice → Voice Design. Describe the voice, generate three previews, save the best match.

What to provide

  • A voice description between 20 and 1000 characters
  • Optional preview text between 100 and 1000 characters
  • Language, gender, age range, tone, accent, and pacing requirements

Details that improve the result

  • Recommended format: Native [Language]. [Gender], [Age]. [Quality]. Persona + Emotion + timbre/pacing sentence
  • Specify language and dialect in the first sentence to prevent drift
  • Avoid FX words like reverb, echo, or phone unless you want degraded quality on purpose

Example prompt

Native English, neutral American accent. Female, 35–40. Studio quality. Persona: trusted product narrator. Emotion: calm, confident, warm. Smooth timbre, relaxed pacing, clear emphasis on key terms.

If the first output is not good

  • Add “studio quality recording” if output sounds thin
  • Replace “accent” with intonation or emphasis when you mean delivery, not dialect
  • Regenerate previews with a shorter, simpler prompt if results feel overfit

Common mistakes

  • Using “accent” when you mean intonation—can trigger unwanted dialect shifts
  • Vague prompts like “nice voice” with no age, pacing, or use case
  • Choosing Voice Design when an existing PVC in the library already fits

Voice remixing

Open a voice you own, use Voice Remixing, describe the transformation in natural language, and compare against the original.

What to provide

  • A cloned or Voice Design voice you personally own
  • The attribute to change: accent, pacing, gender presentation, or audio quality
  • A short test script to validate the remix

Details that improve the result

  • Remixing adjusts attributes while keeping the voice recognizable
  • Useful for character variations, context shifts, or cleaning up clone quality
  • Works with Instant Voice Clones, Professional Voice Clones, and Voice Design voices you own

Example prompt

Remix my existing clone to sound slightly older, slower pacing, and warmer tone for a bedtime story app. Keep the same core identity.

If the first output is not good

  • Change one attribute per remix pass
  • Revert and try a smaller pacing adjustment if identity drifts
  • Test remixed voice on sibilants and plosives, not just one sentence

Common mistakes

  • Stacking many attribute changes in one remix request
  • Remixing voices you do not own
  • Expecting remixing to fix bad source training data

Pronunciation and pacing control

Mark pronunciation inline, then generate. For v3 use /IPA/ transcriptions; for v2 use CMU Arpabet phoneme tags or alias dictionaries.

What to provide

  • Script with hard-to-pronounce names, acronyms, or technical terms
  • Model choice (Eleven v3 supports IPA in forward slashes; v2 supports SSML phoneme tags on Flash v2)
  • Where pauses should fall in the delivery

Details that improve the result

  • Use <break time="1.5s" /> for pauses up to 3 seconds on supported models (not Eleven v3)
  • Eleven v3: wrap IPA in forward slashes, e.g. "/ˌsænfrənˈsɪskoʊ/"
  • Upload pronunciation dictionaries (TXT or PLS) in Studio/Dubbing for recurring terms

Example prompt

The patient takes "/ɡluːˈkoʊs/" daily. <break time="1.0s" /> Next, review the "/ˌdaɪəˈbiːtiːz/" care checklist.

If the first output is not good

  • Swap CMU Arpabet for stubborn v2 words
  • Use alias tags like UN → United Nations for acronyms on Multilingual v2
  • Replace break tags with dashes or ellipses if generations become unstable

Common mistakes

  • Overusing <break> tags in one generation (can cause speed-ups or artifacts)
  • Using phoneme tags on models that do not support them
  • Missing stress markers in multi-syllable phoneme entries

Low-latency and real-time playback

Choose Flash v2.5 for low-latency use cases, keep lines short, and use streaming endpoints for interactive playback.

What to provide

  • Short, conversational lines rather than long narration
  • Flash v2 or Flash v2.5 model selection
  • Streaming use case details (IVR, agent, live demo)

Details that improve the result

  • Flash v2.5 supports 32 languages with a 40,000 character limit and ~75ms ultra-low latency excluding application and network latency
  • Use seed parameter when you need more consistent takes
  • Split very long responses; Flash models prioritize speed over long-form nuance

Example prompt

Model: Flash v2.5. Line: “Got it—checking your order now.” Generate three variants with different seeds and pick the most natural.

If the first output is not good

  • Raise Stability if lines vary too much between regenerations
  • Trim filler from the script to reduce latency further
  • Test on the exact playback device (telephony vs web) and output format

Common mistakes

  • Using Multilingual v2 for sub-100ms interactive loops
  • Assuming identical output every time without a seed
  • Sending entire chat histories in one TTS call

API and developer integration

POST to /v1/text-to-speech/{voice_id} with text and model_id; stream when you need partial playback.

What to provide

  • voice_id from My Voices or the Voice Library (library voices are not available to free-tier API users)
  • Exact text, model_id, and output format (mp3, pcm, opus)
  • Voice settings: stability, similarity_boost, style, use_speaker_boost

Details that improve the result

  • Default model is eleven_multilingual_v2; pick model_id explicitly for Flash or v3
  • Pass previous_text and next_text when chunking long scripts
  • Commercial usage rights require a paid plan even though you retain audio ownership

Example prompt

Convert: “Your table is ready.” voice_id=[narrator], model_id=eleven_flash_v2_5, voice_settings.stability=0.45, similarity_boost=0.75, output_format=mp3_44100_128.

If the first output is not good

  • Log latency per model and region before shipping
  • Cache voice_id + settings pairs that sound best for your product
  • Handle 429s with smaller chunks rather than retrying identical huge payloads

Common mistakes

  • Hard-coding voice IDs you do not control for production workflows
  • Omitting voice_settings and wondering why outputs feel random
  • Assuming Voice Library voices work on free-tier API keys

How to prompt

ElevenLabs speech follows your script and voice choice. Write the line the way it should sound, put emotion in the words (not only in settings), and pick the model for the job—Flash for speed, Multilingual v2 for stable long-form, v3 for expressive performance. For Voice Design, use the structured Native language → gender/age → quality → persona → delivery format from ElevenLabs docs.

Native English, neutral American. Male, 40s. Studio quality.
Persona: SaaS onboarding narrator. Emotion: friendly, confident.
Script: “You are two clicks away from your first export. Let’s walk through it together.”

Put emotion in the script

Models read emotional context from the text. Descriptive cues and punctuation shape delivery more reliably than vague settings alone.

Example

“That’s incredible!” she said, laughing. / “Wait—are you sure about that?” he asked quietly.

Use the Voice Design prompt skeleton

ElevenLabs recommends: Native [Language]. [Gender], [Age]. [Quality]. Persona + Emotion + one or two sentences on timbre and pacing.

Example

Native French, français standard. Female, late 20s. Excellent quality. Persona: airline agent. Emotion: reassuring, precise. Clear articulation, moderate pace, slight smile in tone.

Control consistency with settings

Stability trades emotional range for steadiness; Similarity boost keeps output closer to the source voice. Adjust one at a time.

Example

Narration: stability 0.65, similarity_boost 0.80. Character dialogue: stability 0.35 for wider expression.

Fix pronunciation deliberately

Use IPA slashes on v3, phoneme tags on Flash v2, or alias entries in pronunciation dictionaries for recurring terms.

Example

Our SDK ships with "/kəˈmænd laɪn/" tools and an alias UN → United Nations.

Chunk long content

Split long text into segments for natural prosody. On the API, chain chunks with previous_text and next_text.

Example

Chunk 1 ending: “…and that covers step one.” Chunk 2 starting: “Next, open the dashboard…”

Best output tips

Choose the right model

Flash v2.5 targets ~75ms ultra-low latency excluding application and network latency. Multilingual v2 is the most stable for long-form content. Eleven v3 adds expressive, multi-speaker dialogue with audio tags. Match model to latency, language, and performance needs.

Stage directions get spoken

Text like “she said excitedly” influences delivery but is also spoken aloud. Trim those words in post if you only wanted the emotion, not the narration of it.

Instant vs Professional cloning

Instant Voice Cloning needs roughly 1–3 minutes of audio and returns in seconds. Professional Voice Cloning needs 30–180 minutes, a Creator plan or above, and hours of training—but yields the highest fidelity.

Voice Design when the library misses

Voice Design generates three previews from a 20–1000 character description plus optional preview text. Use it to explore voices; prefer an existing PVC in the library when one already fits.

Explicit language and dialect

Start Voice Design prompts with native language and regional variant. Being explicit reduces accent drift, especially for multilingual projects.

Avoid FX words in Voice Design

Terms like reverb, echo, phone, or tape can degrade quality unless you intentionally want degraded audio. Describe timbre and pacing instead.

Pauses and breaks

Use <break time="x.xs" /> up to 3 seconds on supported models. Eleven v3 does not support SSML break tags—use v3 prompting techniques instead. Too many breaks can destabilize output.

Pronunciation tooling

Eleven v3 accepts IPA wrapped in forward slashes. v2 Flash supports SSML phoneme tags. Studio and Dubbing accept pronunciation dictionaries (TXT or PLS) with alias or phoneme entries.

Long-form prosody

Split very long text and, on the API, pass previous_text and next_text so phrasing flows across chunks instead of resetting every segment.

Consistency and seeds

Outputs are nondeterministic. Use the seed parameter when you need repeatable takes, knowing minor variation may remain.

Free regenerations

You can regenerate the same text with identical voice settings up to two times at no extra cost—useful for occasional glitches, not for script changes.

Commercial rights

You own generated audio, but commercial usage requires a paid subscription. Verify your plan before monetizing outputs.

Voice Library on API

Community Voice Library voices are not available via API to free-tier users. Use voices in My Voices or upgrade accordingly.

Default voice sunset

Save production voice_id values you control (clones, designed voices, or library voices you added). Check ElevenLabs docs for voice availability on your plan.

  • Match voice accent to your script language and target region.
  • Pick the model for the job: Flash v2.5 for latency, Multilingual v2 for long-form stability, Eleven v3 for expressive delivery.
  • Write scripts the way they should sound—emotion lives in the text.
  • Change one variable per revision: voice, stability, script line, or model.
  • For clones, start with clean, consistent source audio before tuning settings.
  • Use Voice Design’s recommended format when library voices do not fit.
  • Split long narration into segments instead of one massive generation.
  • Check plan limits for Voice Library API access and commercial usage rights.
  • Regenerate identical text up to twice for free if you hit a rare quality glitch.

Try this AI

Try ElevenLabs

Product Details

Pricing, features, limits and latest updates

ElevenLabs

AI voice generator and voice agents platform for lifelike speech, cloning, and conversational agents.

Free / Paid · Free

Pricing Plans

Free

Free

Starter

$6/ Monthly

Starter

$60/ Yearly

Creator

$22/ Monthly

Creator

$220/ Yearly

Pro

$99/ Monthly

Pro

$990/ Yearly

Scale

$299/ Monthly

Scale

$2,990/ Yearly

Business

$990/ Monthly

Business

$9,900/ Yearly

Enterprise

Custom

Key Features

Voice

Shared monthly credits cover text to speech, dubbing, speech to text, and sound effects.

API

API usage is billed in US dollars per model, not from Creative plan credits.

Agents

ElevenAgents is a separate conversational-agent product with its own pricing surface.

Music Generation

Eleven Music draws from the same shared monthly credit pool as speech products.

Voice Cloning

Instant and professional voice cloning are plan-gated features on the Creative subscriptions.

Limits

  • Eleven v3 request limit: 5000 characters. Maximum characters per Eleven v3 TTS request on the API (ElevenLabs models docs).
  • Eleven v3 SSML: SSML break tags are not supported on Eleven v3; use audio tags and punctuation instead (ElevenLabs best practices).

Ideas / Prompt experiences

Share a prompt that worked for you. Username and email are shown with your submission. External links are not allowed.

Example prompt

TTS narration with delivery tags (v3)

Prompt

[calm] Welcome back. In this lesson, we'll configure webhooks so your app receives events in near real time. [pause] Let's start with the payload shape.

Short explanation

Eleven v3 supports audio tags such as [calm] and [pause]; SSML break tags are not supported on v3 per official docs.

Example prompt

Voice selection for product demo

Prompt

Voice: clear female narrator, mid-30s, neutral accent. Pace: measured. Length: ~30 seconds. Script: Explain why async jobs beat synchronous requests for file processing.

Variation

Generate three variants with different seeds and pick the most natural take.

Short explanation

Specify voice, pace, and length before the script; TTS consumes credits per character on paid plans.

Share your experience

Required fields are marked. Variation, result, and explanation are optional.

Suno logo

Suno

Explore

Music generation from a prompt, including vocals and full songs.

AudioFree / Paid

ACE Studio logo

ACE Studio

Explore

AI music studio for singing vocals, instruments, voice cloning, and stem splitting.

AudioPaid

AirMusic logo

AirMusic

Explore

Turns a text prompt or lyrics into an original song, with optional vocals and a music video.

AudioFree / Paid