ElevenLabs is growing its revenue at roughly double the pace most AI startups manage, and now it’s dropping a model update that actually matches that ambition. The company launched two new speech models on Monday, ElevenLabs v4 and v4 Turbo, with meaningful improvements across expression control, voice cloning speed, language coverage, and latency for real-time voice agents. This isn’t a cosmetic refresh.
The headline feature for most developers will be cloning. V4 can clone a voice from just 10 seconds of audio, which is a real reduction from what earlier versions required. The new architecture also handles voice identity better across longer text passages, something that has historically been a weak spot for AI narration tools. And with expanded inline tags, users can now stack multiple expression cues in sequence, giving much finer control over how a line actually sounds. ElevenLabs introduced these tags in v3, but v4 makes them genuinely composable.
Language support jumped from 70 to 90, and the company says the biggest quality gains came in Japanese, Brazilian Portuguese, Mandarin, and Cantonese. That’s not random. ElevenLabs has been hiring aggressively in India, Europe, and Brazil, and those markets require this kind of multilingual depth to compete locally. The company’s headcount has crossed 800, and over 55% of its revenue now comes from enterprise clients.
The v4 Turbo variant is aimed squarely at voice agents, where latency is the thing that makes or breaks the product. The model can start generating audio as soon as the underlying LLM begins producing a response, rather than waiting for a full output. It also handles holds, escalations, and confrontational exchanges differently, which matters if you’re deploying agents in customer service or sales contexts. This puts ElevenLabs in direct competition not just with startups like Cartesia, Deepgram, Fish Audio, and WellSaid Labs, but also with Google and OpenAI, both of which have been improving their own voice capabilities.
The competitive framing matters here. The speech model market has filled up quickly, and differentiation is getting harder. What ElevenLabs has that most rivals don’t is scale: an annualized revenue run rate that has climbed from around $330 million at the start of the year to over $600 million. The company raised $500 million from Sequoia at an $11 billion valuation earlier this year, and there are already reports of a follow-up round that could push that number to $22 billion.
CEO Mati Staniszewski has said an IPO is on the horizon, though without a firm timeline. V4 is the kind of release that keeps that story credible.



