ElevenLabs has launched Eleven v4 for narration and character performances and Turbo for low-latency voice agents, AlphaSignal reports. The release adds expression controls and broader language support, while v4’s rebuilt cloning system can create a voice from a 10-second recording and brings back professional cloning.
Expression controls and phonetic pronunciation
According to AlphaSignal, the two variants divide speech generation by workload: Eleven v4 focuses on narrated material and character delivery, while Turbo prioritizes low-latency voice agents. Both include controls for vocal expression, adding a way to shape the delivery of generated speech alongside their different performance priorities.
AlphaSignal reports support for more than 90 languages in both variants, compared with approximately 70 in v3. Eleven v4 also adds IPA phonetic control, allowing pronunciation to be specified phonetically rather than relying solely on the written text. That pronunciation control sits alongside the expression controls shared by the two variants.
Pronunciation improves, with uneven category results
AlphaSignal reports that Eleven v4 scored 91.7% on Artificial Analysis’ composite pronunciation-robustness benchmark, compared with 85.6% for v3—the highest result measured in that benchmark at the time. The evaluation uses human reviewers who compare generated audio with agreed pronunciations, testing challenges that include abbreviations, technical terms, context-dependent words and character sequences.
The category results show where the improvement is concentrated. According to AlphaSignal, v4 scored 94.1% for correctly expanding abbreviations, against v3’s 82.7%. It also scored 94.1% when selecting pronunciation from sentence context, although Gemini 3.8 Flash led that category with 97.9%. For standalone terms without sentence context, v4 reached 93.2%, behind v3 Conversational’s leading 95.1%.
Exact identifiers and character strings remained v4’s weakest category, AlphaSignal reports. Its 78.8% result improved on v3’s 71.2%, but SpaceXAI TTS led with 85.7%. This category tests preservation of the precise sequence being read, rather than abbreviation expansion or pronunciation selected from the surrounding sentence.
English listener rankings depend on the voice setup
At the time of AlphaSignal’s article, Eleven v4 led Artificial Analysis’ English Provider Voice Arena with 1,319 Elo, ahead of Cartesia Sonic 3.6 at 1,276 and Gemini 3.8 Flash TTS at 1,267. Its result drew on 1,674 evaluation samples. This arena compares listener preferences using each model’s own voices and default configuration, with higher Elo indicating more frequent preference in head-to-head evaluations.
Using a common voice produced a different leader. AlphaSignal reports that the Controlled Voice board, which uses the same cloned voice across systems to reduce the influence of providers’ voice catalogs, placed Eleven v4 second at 1,157 Elo. Alibaba’s Qwen-Audio-3.1-TTS-Plus led with 1,178.
The English result also has a specific language scope. According to AlphaSignal, Cartesia Sonic 3.6 continued to lead eight of the nine non-English language boards tracked by Artificial Analysis. Those language-specific standings differ from the English Provider Voice Arena result, despite v4’s support for more than 90 languages.
Faster synthesis and Turbo streaming
Artificial Analysis measured Eleven v4 generating 73.4 characters per second, up from 42.5 for v3, according to AlphaSignal. This is synthesis throughput: the measurement excludes network delays, the time a language model takes to generate its response and audio playback on the client.
For Turbo, AlphaSignal reports a median time to first speech of 150 milliseconds and support for bidirectional streaming. The interval runs from a synthesis request to the first returned audio chunk. A voice agent’s complete response timing also includes network conditions, language-model generation, buffering and playback.
Existing voice clones need retraining
AlphaSignal describes a rebuilt cloning system for Eleven v4. A recording lasting 10 seconds is enough for an instant voice clone. The professional cloning option, absent from v3, is available again. To work effectively with v4, older clones need to be trained again, whether they were produced through the instant or professional process.
Applications already using ElevenLabs’ conversion API can select v4 by changing the model identifier, according to AlphaSignal. That application-side selection and the preparation of existing voices are separate parts of migration: changing the identifier does not replace the required clone-retraining step.
API and Creator+ access
According to AlphaSignal, the standard API price for v4 is $80 per million characters. That is approximately 1.6 times the price of Sonic 3.6 and 4.9 times that of Gemini 3.8 Flash TTS. A launch offer lasting two weeks lowers the price to $22 for v4 and $11 for Turbo per million characters. Actual prices can vary with the plan, volume agreement and included credits. Eligible Creator+ subscribers can use v4 within their monthly credit allowance without paying an extra model surcharge.



