ElevenLabs launched two new speech models on Monday, called ElevenLabs v4 and v4 Turbo, offering more expression control, lower latency for voice agents, and support for more than 90 languages.
The company released its v3 model last year, and teased the model at an event in Warsaw earlier this year. For the v4 generation of models, ElevenLabs is adopting a new architecture that allows for better control and faster cloning. The company said that with v4, users will be able to clone a voice with just 10 seconds of audio.
ElevenLabs introduced inline tags to define expression with v3, and is expanding those tags in v4, letting users stack multiple tags and having the model follow the sequence.
The previous version supported 70 languages, and ElevenLabs has worked to get that number up to 90 languages with the new version. The startup said that it observed the biggest quality jump in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
The company said that the new model is suited for voice agents, as the new version has lower latency to allow for more fluid conversation.
Competition in speech models has ramped up as startups like Cartesia, Deepgram, Fish Audio, Boson, and WellSaid Labs have created expressive speech models.
ElevenLabs raised $500 million earlier this year from Sequoia earlier this year, in a round led by Sequoia that valued the company at $11 billion.