
Eleven v4 and v4 Turbo
ElevenLabs' new voice generation: script-aware delivery for produced audio, 100 ms latency for live agents.
Quick verdict
- Best for
- Audiobook, podcast, and video narration with multiple characters
- Pricing
- Included in ElevenLabs plans; free tier with limited characters, paid from $5/mo
- Not ideal if
- Latency figures are the vendor's own
What is Eleven v4 and v4 Turbo?
Released on September 28, 2026, Eleven v4 is ElevenLabs' quality-focused text-to-speech model on a new architecture that reads a script with awareness of who is speaking and what came before, supports more than 90 languages, multiple speakers, and inline directions such as [laughs] or [door slams]. Eleven v4 Turbo targets live calls and agent loops with about 100 ms median inference latency and 150 ms to first speech. Both are available in ElevenLabs products and the API; Turbo also ships through ElevenAgents.
Key features
- Script-aware, multi-speaker delivery
- Inline direction tags for emotion and sound effects
- 90+ languages
- Turbo: ~100 ms median latency, ~150 ms to first speech (ElevenLabs-reported)
- Available via API and ElevenAgents
Pros
- Directions in the script replace fiddly per-line settings
- Two models cover both produced audio and real-time agents
- Available to existing subscribers at no new price
Cons
- Latency figures are the vendor's own
- Character-based pricing still adds up for long content
- Voice cloning rules and consent checks apply as before
Best for
Read more
Related comparisons
Alternatives to Eleven v4 and v4 Turbo
ElevenLabs
The AI voice generator with the most realistic output.
Cartesia Sonic
The fastest realistic AI voice — 90 ms latency, indistinguishable from human.
Speechmatics
Speech-to-text API known for accuracy across accents and 50+ languages.


