ElevenLabs released a new text to speech model on Monday that it says can carry emotional tone, not just read words aloud. The company calls Eleven v4 its most expressive voice model yet, alongside a faster version built for live conversations, Eleven v4 Turbo.

The stakes here are less about a single model and more about a stubborn tradeoff in voice AI: expressive voices have tended to be slow, and fast voices have tended to sound flat. A customer service bot that speaks quickly but monotone is a familiar and mildly frustrating experience for anyone who has called a support line in the past few years.

ElevenLabs says Eleven v4 topped Artificial Analysis’ Provider Voice Arena leaderboard for September 2026. The company also reports that graders who judged the model against rivals without knowing which was which, ties counted as half a vote, picked Eleven v4 roughly 75 percent of the time. Both figures come from the company’s own release and its cited testing methodology, not from an independent audit.

The faster variant, Eleven v4 Turbo, has a median time to first audible speech of about 150 milliseconds, according to ElevenLabs, which it measured against Cartesia’s Sonic 3.6, xAI’s TTS, Google’s Gemini 3.8 Flash-Lite TTS, and OpenAI’s GPT-4o mini TTS. The company says that speed is close to the average pause length in real human conversation, which is the practical bar for making a voice agent feel responsive rather than laggy.

Beyond speed, ElevenLabs is emphasizing control. Users can now describe delivery in plain language or insert tags like “[laughs],” “[said angrily in French accent],” or “[phone buzzing]” directly into a script, and the company says the model follows those instructions more reliably than earlier versions. IPA phoneme support, used for custom pronunciations, has also been improved.

Voice cloning gets an upgrade too. ElevenLabs says its Instant Voice Clone feature can now produce a usable clone from just 10 seconds of source audio, and that cloned voices hold their character more consistently across long projects like audiobooks or ad campaigns. The company also added Professional Voice Clone support to the new models for higher-fidelity use cases. Both Eleven v4 and Eleven v4 Turbo cover more than 90 languages, and ElevenLabs says accent drift, where a cloned voice slowly reverts to its original accent during a long generation, is now noticeably reduced.

Both models ship built for ElevenLabs’ own agent stack, ElevenAgents, rather than as components meant to be assembled with outside vendors’ tooling. The company frames that vertical integration as an advantage over competitors who stitch together speech and orchestration layers from different providers, though it offers no head to head latency comparison for that specific claim. The company has not disclosed pricing changes alongside the release.

For any team currently running a voice agent on a model tuned for speed over expressiveness, this is the moment to re-test call transcripts against Eleven v4 Turbo, particularly in healthcare, retail, or gaming support lines where a flat tone measurably hurts customer satisfaction scores.

Reported by ElevenLabs in its own product announcement on 28 September 2026.