ElevenLabs released v4, a speech model that follows direction cues more accurately and maintains voice consistency across long productions. It is the company's first major update since v3, launched just over a year ago.

Eleven v4 scores 91.7 percent on the pronunciation benchmark, up from v3's 85.6 percent. That compares with 85.6 percent for its predecessor, according to the company.

The model is built on a new architecture and targets real-time voice agents. Availability begins with temporary price cuts, initially for developers and creators.

"In Elevenlabs' tests, Turbo starts producing audible speech in 150 milliseconds, compared with 262 milliseconds for Cartesia Sonic 3.6," said Jonathan Kemper, the company's spokesperson. The model also outperforms OpenAI's GPT-4o mini TTS, which takes 814 milliseconds.

The announcement follows the release of Eleven v3, which already supported audio tags but followed them less accurately. ElevenLabs said the update improves upon that by making voices more consistent and expressive.

ElevenLabs did not say how long the price cuts will last, and it raised questions about data storage and privacy. The company said it will store customer data in the US by default, though enterprise customers can store data in isolated environments in the EU, India, or Singapore.

Source: thedecoder