Quick answer
ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28, 2026. v4 is the quality model for produced audio: built on a new architecture, it reads a script with awareness of who is speaking, what happened before, and how each line should land, supports more than 90 languages, multiple speakers, sound effects, and inline directions such as [laughs], [whispers], and [door slams]. v4 Turbo targets live calls and agent loops with about 100 ms median inference latency and 150 ms median time to first speech, figures ElevenLabs reported. Both are available in ElevenLabs products and the API; Turbo also ships through ElevenAgents.
Text-to-speech has had two different customers for a while: people producing audiobooks, videos, and podcasts, who care about performance, and people building voice agents, who care about the gap between the caller finishing a sentence and the agent starting one. v4 is the first ElevenLabs generation that gives each customer their own model.
What v4 does differently
- Context awareness: the model reads the whole script, so a line is delivered knowing who said the previous one and what the scene is
- Directions in the text: [laughs], [whispers], [sighs], [door slams] and similar tags replace fiddly per-line settings
- Multiple speakers in one script, with consistent voices
- More than 90 languages, with the emotional range carried across them
What Turbo is for
A voice agent feels natural when it starts speaking within a few hundred milliseconds of the caller stopping. Turbo's reported 100 ms median inference and 150 ms to first speech, combined with the speech-recognition and model-response steps, is the budget that makes that possible. It is offered through ElevenAgents as well as the API, which is where most production voice agents on ElevenLabs are built. As always with latency, measure it from your own region and stack; vendor medians are measured in ideal conditions.
Who should switch
- Narration with several characters or emotional range: v4, immediately
- Live agents on ElevenLabs: Turbo, after testing interruption handling and real-world latency
- Simple single-voice announcements: the previous models remain fine and may be cheaper per character
- Anyone cloning a real voice: the consent and verification rules are unchanged
Bottom line
v4 makes directed, multi-character audio far easier to produce, and Turbo puts ElevenLabs squarely in the low-latency race it had been losing to specialists. If you already pay for ElevenLabs, both are included; try v4 on a scene with two characters and you will hear the difference.

