Working with TTS to produce multi-character audio used to be pure manual labor: generate each character one at a time, align the timelines yourself, insert pauses, and pray the synthesized emotions didn't sound too robotic. ElevenLabs V3 changes exactly that workflow — it supports generating dialogue for multiple speakers within a single request, letting the model handle turn-taking, emotional continuity, and natural pauses instead of leaving that work to you.
It's now available to developers in Alpha through the APIPASS platform.
What ElevenLabs V3 Actually Does
The most fundamental shift in V3 is the move from "synthesizing speech" to "generating dialogue." Previous models handled individual voice lines; V3 handles an entire multi-person conversation — it understands context, and knows that after Character A delivers an angry line, Character B's response should carry a specific tone in return.
Things that weren't possible before are now on the table:
- Multi-character dialogue, completed in a single request: No manual audio splicing, no timeline alignment. The model natively supports multiple speakers and automatically handles turn-taking, interruptions, and overlapping speech.
- Emotion stays coherent across the conversation: If Character A is angry, Character B's response won't suddenly shift into a flat broadcast tone. V3 understands the emotional context of a conversation.
- Audio tag instructions control performance details: Through inline tags like
[whispers],[laughs], and[interrupting], you can write directly into the script that "this line should be low and quiet" or "this character cuts the other one off" — no post-processing required. - Support for 70+ languages: Emotional expressiveness carries over into languages beyond English.
Audio Tags: Controlling Performance Through Text
V3 introduces an "Audio Tags" system that lets you embed instructions directly in text using square brackets. These tags tell the model how to perform a line, not just read it.
| Type | Examples |
|---|---|
| Emotion | [excited], [sad], [angry], [flustered] |
| Manner of Speaking | [whispers], [shouts], [dramatic tone], [sings] |
| Natural Reactions | [laughs], [sighs], [clears throat], [light chuckle] |
| Style / Context | [French accent], [auctioneer], [pirate voice] |
| Sound Effects | [applause], [gunshot], [leaves rustling] |
| Dialogue Dynamics | [interrupting], [overlapping], [cuts in] |
Tags can be stacked — for example, [French accent] [whispers] makes a character speak in a low, French-accented voice. In V3's demo scripts, the [sings] tag even prompted the model to sing a birthday song — V3 isn't just "rhythmic speech," it genuinely enters melodic territory.
ElevenLabs V3 vs. Eleven Multilingual v2: How to Choose
| ElevenLabs V3 (Alpha) | Eleven Multilingual v2 | |
|---|---|---|
| Positioning | Expressive, dialogue generation | Stable, long-form content |
| Language Support | 70+ languages | 29 languages |
| Multi-Speaker | Natively supported (single request) | Manual line-by-line splicing required |
| Audio Tags | Full support, including sound effects | Basic support (pauses, phrasing) |
| Max Characters per Request | 5,000 | 10,000 |
| Best For | Games, audiobooks, interactive media | Enterprise voiceover, e-learning, long-form content |
A note on latency: If you're building a real-time voice agent, V3 is not your best option — Eleven Flash v2.5 (approximately 75ms) is better suited for that. V3 is optimized for high-fidelity performance, not low-latency response.
What ElevenLabs V3 on APIPASS Is
APIPASS is a third-party API aggregation platform that wraps multiple AI services under a unified interface and authentication system. The official ElevenLabs API comes with application requirements and quota management; through APIPASS, developers can call ElevenLabs V3 using the same API key and request format without needing to register a separate ElevenLabs account or manage its quota logic independently.
APIPASS uses an asynchronous task model for ElevenLabs V3: you first send a request to create a task and receive a taskId, then poll to check the task status, and retrieve the audio file URL once it's complete. This differs from the ElevenLabs official API's synchronous response pattern — it reflects APIPASS's unified task queue architecture.
Understand ElevenLabs V3 API Parameters
The APIPASS Playground gives you a direct, no-code view of every parameter the ElevenLabs V3 API accepts. Here's what each field means and how it maps to your actual API request.
dialogue (Required)
The dialogue field is the core of every request. It takes an array of conversation entries, and each entry must include two things: a text field and a voice field. The Playground lets you add as many dialogue entries as needed using the + Add dialogue button, and remove any entry with Remove.
Each dialogue entry works independently — a different voice can be assigned per entry, and Audio Tags can be embedded inline within the text itself. The two entries shown in the Playground illustrate this well:
- Dialogue 1:
[applause] Thank you all for coming tonight! Today we have a very special guest with us.— Voice: British Football Announcer - Dialogue 2:
[gulps] ... [strong canadian accent] [excited] Hello everyone! Thank you all for having me tonight on this special day.— Voice: British Football Announcer
Notice that Dialogue 2 stacks three tags in sequence — [gulps], [strong canadian accent], and [excited] — demonstrating how the tag system handles both physical reactions and stylistic modifiers within the same line.
stability (Optional)
The stability parameter controls how consistent and predictable the voice performance is. The adapter accepts only three discrete values:
| Value | Behavior |
|---|---|
0 | Maximum expressiveness; higher variability between generations |
0.5 | Balanced default; used automatically when the field is omitted |
1 | Maximum stability; most consistent but least expressive |
The Playground defaults this to 0.5, which is the recommended starting point for most use cases. If you're generating final production audio and need reproducible results, move toward 1. If you're generating character voices where emotional range matters more than consistency, try 0.
language_code (Optional)
The language_code field accepts an ISO 639-3 three-letter language code — for example, eng for English, fra for French, or jpn for Japanese. This value is forwarded directly to the ElevenLabs upstream provider without modification.
Leaving this field empty instructs the provider to detect the language automatically from the input text. For most single-language scripts, automatic detection works reliably. If your dialogue mixes languages or uses an accent tag like [strong canadian accent], specifying the base language code explicitly can help anchor the model's pronunciation expectations.
Output
Once you hit Run, the Output panel on the right renders the generated audio directly in the browser via a standard audio player. The output type is audio, and the result is an MP3 file you can preview, download, or retrieve programmatically via the resultUrls field in the task query response.
The View full history button lets you revisit previous generations within the same session — useful when you're comparing multiple stability settings or tag combinations side by side.
How to Call the ElevenLabs V3 API
The entire flow is two steps: Create Task and Query Task.
Step 1: Create a Task
Send a POST request to /api/v1/jobs/createTask with your dialogue content in the body.
Request Example:
curl -X POST "https://api.apipass.dev/api/v1/jobs/createTask" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "elevenlabs/text-to-dialogue-v3",
"input": {
"dialogue": [
{
"text": "[excitedly] Hey Jessica! Have you tried the new ElevenLabs V3?",
"voice": "Liam"
},
{
"text": "[whispering] Yeah, just got it! The emotion is so amazing. [laughs]",
"voice": "Jessica"
}
],
"stability": 0.5,
"language_code": "eng"
}
}'
Key Parameter Reference:
| Parameter | Type | Description |
|---|---|---|
model | string (required) | Fixed value: elevenlabs/text-to-dialogue-v3 |
input.dialogue | array (required) | Dialogue array; each entry contains text and voice. Total character count must not exceed 5,000 |
input.stability | number (optional) | Accepts 0, 0.5, or 1. 0 is most expressive but most variable; 1 is most stable. Defaults to 0.5 |
input.language_code | string (optional) | ISO 639-3 language code, e.g. eng, fra, jpn. Leave blank for automatic detection |
callbackUrl | string (optional) | Callback URL triggered when the task completes |
Response Example:
{
"code": 200,
"message": "success",
"data": {
"taskId": "task_elevenlabs_dialogue_1768468200000"
}
}
Once you have the taskId, move to Step 2.
Step 2: Query the Task
Send a GET request to /api/v1/jobs/recordInfo with the taskId to check the status.
Request Example:
curl -X GET "https://api.apipass.dev/api/v1/jobs/recordInfo?taskId=YOUR_TASK_ID" \
-H "Authorization: Bearer YOUR_API_KEY"
Response Example (Success):
{
"code": 200,
"message": "success",
"data": {
"taskId": "task_elevenlabs_dialogue_1768468200000",
"model": "elevenlabs/text-to-dialogue-v3",
"state": "success",
"resultJson": {
"resultUrls": [
"https://cdn.apipass.dev/apipass/results/task_elevenlabs_dialogue_...mp3"
]
},
"failCode": null,
"failMsg": null,
"costTime": 0,
"completeTime": 1768468205000,
"createTime": 1768468200000,
"modelRequested": "elevenlabs/text-to-dialogue-v3",
"modelUsed": "elevenlabs/text-to-dialogue-v3",
"fallbackTriggered": false
}
}
Key Response Fields:
| Field | Description |
|---|---|
data.state | Task status: queuing, processing, success, or fail |
data.resultJson.resultUrls | Array of generated MP3 file URLs |
data.failCode / data.failMsg | Error code and message when the task fails |
data.costTime | Processing time elapsed |
data.completeTime | Unix timestamp of task completion |
Tasks typically require several polling rounds. A 2–3 second interval between polls is recommended until state returns success or fail.
Usage Recommendations
Keep individual inputs short: V3's limit is 5,000 characters, but staying under 2,000 characters produces more consistent results. With very long scripts, a character's emotional state can "drift" — the angry tone established at the start may flatten into bland narration by the end.
Use punctuation to reinforce delivery:
- Ellipses (
...) create natural trailing-off, useful for hesitation or thoughtful pauses. - Em dashes (
—) signal an interruption. Place a—at the end of Character A's line and the beginning of Character B's, and the model will simulate a natural mid-sentence cut-in.
Don't stack too many tags: A single [sad] at the start of a passage works better than tagging every sentence. Too many tags, and the speech starts to sound forced rather than natural.
Generate multiple takes and pick the best: V3 has inherent variability — identical inputs won't always produce identical results. For production content, generate 2–3 versions and select the strongest one.
Voice selection: During the Alpha phase, Instant Voice Clone (IVC) or the platform's built-in Designed Voices are the recommended options. Professional Voice Clone (PVC) is still being optimized for the V3 architecture, and clone quality may fall short of expectations.
Use Cases
Indie game development: This is where V3's practical value is most apparent. Independent developers who previously needed NPC dialogue audio had two options: hire voice actors (expensive) or settle for flat TTS (poor player experience). V3's multi-character dialogue and interruption mechanics make NPC conversations sound like real people talking, not characters taking turns reading from a script. The cost of producing audio for dialogue systems drops significantly while content density increases — budgets that previously only covered main story lines can now extend to side quests and ambient NPC chatter.
Audiobooks and podcasts: One writer, one request, a fully generated multi-character conversation. The narrator shifts tone, characters cut each other off, laughter and sighs land naturally — all of it previously required manual post-production work, and all of it can now be resolved at the generation stage.
Multilingual content localization: Brands publishing video content across regions need localized voiceover in multiple languages. V3 carries full emotional expression across 70+ languages — not generic machine-translation delivery, but voiceover that fits local spoken idioms.
Getting Started
- Register an account on APIPASS and access the developer console.
- Generate an API Key to authenticate your
createTaskandrecordInforequests. - Test in the APIPASS ElevenLabs V3 Playground — you can fill in dialogue content, select voices, and adjust parameters directly in the interface. Results play back in real time on the right side, no code required.
- Once you're satisfied with the output, integrate using the curl examples above into your backend.
