The Decoder reports that ElevenLabs is releasing Eleven v4, a speech model the company says follows direction cues more accurately and keeps voices consistent across long productions. A new architecture also powers the Turbo variant for real-time voice agents. Eleven v4 generates laughter, whispers, and sounds such as slamming doors more reliably than its predecessor. Turbo starts producing speech in about 150 milliseconds. Eleven v3, released just over a year ago, already supported these audio tags, but followed them less accurately.
[1]
ElevenLabs says v4 uses a new architecture that analyzes a script’s tone, pacing, and context. Users can give directions through tags or plain sentences, and they can use phonetic spelling for names and technical terms. The company says pronunciation controls now work more reliably. Narrators and characters should stay consistent through a production, even when a line is regenerated several times. One request handles up to 10,000 characters, roughly ten minutes of audio. Longer works such as audiobooks use multiple segments, with pacing and delivery expected to hold across the joins. In dialogue, speakers respond to the whole scene rather than to each line alone. The new version supports more than 90 languages, up from about 70 in v3. Cloned voices should speak other languages with a local accent, without drifting back to the original accent over time. Professional Voice Clones work again after being unavailable in v3. An Instant Voice Clone needs only ten seconds of audio.
[1]Turbo is for real-time uses such as customer-service calls or game characters. The company says voice-agent developers previously had to choose between speed and expression, and that v4 Turbo is meant to offer both. In ElevenLabs’ own tests, Turbo starts producing audible speech in 150 milliseconds, Cartesia Sonic 3.6 in 262 milliseconds, and OpenAI’s GPT-4o mini TTS in 814 milliseconds. ElevenLabs says it optimized Turbo together with the ElevenAgents platform. Those are the company’s tests, not a repeat measurement here. The Decoder also writes that Eleven v4 ranks ahead of Cartesia Sonic 3.6 and Google’s Gemini 3.8 Flash TTS on Artificial Analysis’ Provider Voice Arena leaderboard. It scores 91.7 percent on the pronunciation benchmark, up from 85.6 percent for v3. In ElevenLabs’ blind tests, about three-quarters of listeners preferred v4 over models from Cartesia, Inworld, and Google. In those blind tests, listeners rated v4 as more expressive in 65 to 81 percent of comparisons, depending on the competitor.
[1]The company points to higher-quality clones for dubbing: an actor’s voice used across every supported language. ElevenLabs licenses voices from the people behind them. A voice actor can offer a heavily trained clone in the library and earn money when paying users use it. Celebrity voices, including Michael Caine’s, are already offered through a dedicated marketplace. Both models launch with temporary price cuts. Standard API pricing is $80 per million characters for v4 and $40 for Turbo. Through October 12 those rates fall to $22 and $11. According to ElevenLabs, users on the $22 monthly Creator plan or higher can use v4 in ElevenCreative at no extra cost for two weeks, capped at twice their monthly credits. Artificial Analysis lists Sonic 3.6 at $49 per million characters and Gemini 3.8 Flash TTS at $16.49. Both models are available now in ElevenAgents, ElevenCreative, and through the API. Documentation says customer data is stored in the US by default. Enterprise customers can store it in isolated environments in the EU, India, or Singapore, though some processing may take place outside the chosen region. In the EU, a mode that does not retain data can keep API processing inside the region.
[1]要点
- ElevenLabs says v4 follows direction cues more accurately and stays consistent in long productions, up to about 10,000 characters per request.
- Language coverage rises from about 70 to more than 90. Professional Voice Clones return. An instant clone needs ten seconds of audio.
- In the company’s tests Turbo starts speaking in 150 ms, against 262 ms for Sonic 3.6 and 814 ms for GPT-4o mini TTS.
- API list prices are $80 and $40 per million characters, cut to $22 and $11 through October 12. Data sits in the US by default.