ElevenLabs Releases v4 Speech Models With More Voice Control and 90-Plus Languages
The v4 release also targets quicker voice cloning and less waiting in AI calls, though its speed and quality improvements remain company claims.
Loading page…
The v4 release also targets quicker voice cloning and less waiting in AI calls, though its speed and quality improvements remain company claims.
Listen to this story
ElevenLabs’ v4 release targets both production control and live-agent performance: stacked in-script direction is meant to shape delivery, while audio can begin streaming before an agent’s full response is generated. The models expand supported languages from 70 to more than 90, but ElevenLabs reports the strongest quality gains only in Japanese, Brazilian Portuguese, Mandarin and Cantonese. Voice cloning from 10...
Compared with v3, users can stack expression tags and direct how the model follows their sequence.
The system is designed to preserve a recognizable voice across longer passages while using context to adjust delivery.
ElevenLabs says v4 can vary delivery for confrontations, escalations and holds—claims that need validation in changing live conversations.
ElevenLabs has released v4 and v4 Turbo, two speech models built to give creators finer control over how words sound and help voice agents answer with less delay. The models support more than 90 languages, up from 70 in the previous version. For businesses putting AI on calls, the key question is whether the company’s promised speed and expression hold up in a real conversation.
ElevenLabs already let users place expression tags inside a script with v3. In v4, users can stack those tags, and the model is designed to follow their sequence. That gives someone producing a narrated scene more direction over changes in delivery than a single instruction can provide.
The model is also designed to keep a voice recognizable through longer passages while using the surrounding text to guide expression. Those are different jobs: preserving who appears to be speaking and adjusting how that speaker delivers a line. ElevenLabs says a new model architecture improves its control and speeds up cloning, but the reported capabilities do not by themselves show how consistently a finished recording will sound.
The larger language count is a coverage change, not a promise that every language improved equally. ElevenLabs says it saw its biggest quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. That distinction matters to anyone making speech for several markets: the list of available languages says where a model can be used, while the quality of the resulting voice remains a separate judgment.
For a voice agent, expression is only part of the problem. A caller also notices the silence before an answer starts. ElevenLabs says v4 has lower latency and can begin making audio while the language model behind an agent is still generating its reply. Rather than waiting for the entire answer to be written, the speech system can start speaking as text arrives.
The company also says the model can deliver confrontations, escalations and holds differently to help resolve issues. That is a more demanding claim than simply producing a natural-sounding sample. On a call, the delivery has to fit a changing exchange, and a quick start matters only if the words and tone still make sense as the answer unfolds.
ElevenLabs says v4 can clone a voice from just 10 seconds of audio. If that works as described, the short recording could make it easier to set up a voice for a project. It also makes permission an important practical question: the amount of audio needed to copy a voice says nothing about whether the person behind it has agreed to that use.
Large companies already account for more than 55% of ElevenLabs’ business, according to the company. That helps explain why this release pairs tools for directed speech with a pitch for smoother automated calls. For those customers, the useful test is not a language count or a short voice sample alone. It is whether the model can keep its delivery convincing and its responses timely across the calls they actually handle.
Loading discussion...
Join the conversation
What would make those checks reasonable?
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.