Voice AI Cluster on Sep 30: ElevenLabs at $22B, Inworld Buys Ultravox, Eleven v4 Hits Sub-150ms Turbo
Three concrete voice-AI moves on Sep 30: ElevenLabs employee tender at $22B valuation ($300M, co-led by Wellington and T. Rowe Price), Inworld acquires Ultravox voice-agent platform, ElevenLabs Eleven v4 + Turbo model hits sub-150ms latency.
Sep 30, 2026 was a clustering day for voice AI. Three concrete moves landed within 24 hours: ElevenLabs doubled its valuation to $22 billion in a $300 million employee tender offer, Inworld acquired the Ultravox voice-agent platform, and ElevenLabs shipped Eleven v4 with a Turbo variant that hits sub-150ms latency. None of these were predicted by pre-event leaks — they came in through separate channels and the cluster only reads as a cluster in retrospect.
ElevenLabs at $22B: $300M tender, Wellington + T. Rowe Price co-led
The headline number is the valuation. ElevenLabs is now valued at $22 billion in a $300 million employee tender offer, co-led by Wellington and T. Rowe Price. The tender doubled the company’s $11 billion valuation from its February 2026 primary raise ($500M at the time).
This is the second tender ElevenLabs has authorized. The first was a $100M tender at a $6.6B valuation in September 2025. The pattern is consistent with what TechCrunch’s earlier coverage called the “secondary sales shift from founder windfalls to employee retention tools” — large AI startups use periodic employee liquidity as a retention mechanism to prevent staff from leaving for competitors. ElevenLabs is doing exactly that, with the second tender 12 months after the first, and at 3.3x the per-employee valuation.
Four-year-old ElevenLabs (founded 2022) is now New York + London-based and joins the ranks of Europe’s most valuable startups. Co-founder/CEO Mati Staniszewski is the public face; TechCrunch sat down with him last week for context.
For builders, the meaningful detail isn’t the $22B headline — it’s that the tender is structured for institutional retention (Wellington and T. Rowe Price are the kinds of investors who hold stock through an IPO). That’s a signal ElevenLabs is positioning for public-market readiness within the next 12-24 months, not just continued private growth.
Inworld acquires Ultravox
Inworld announced on September 30 that it acquired Ultravox, the real-time voice-agent platform. Financial terms were not disclosed. Per Inworld’s announcement, Ultravox’s existing customers will see continuity of third-party TTS providers — Inworld’s own built-in voices move to Realtime TTS-2 at no extra cost and with no code changes.
Ultravox’s product positioning is “open-weight speech-to-speech model” — different from text-to-speech (ElevenLabs, Inworld’s own stack) and different from cascaded ASR + LLM + TTS (the older voice-agent architecture). The acquisition is a vertical integration play: Inworld gets both the open-weight model and the developer platform in one move, with the developer platform’s existing customer base and toolset as the immediate asset.
For builders running voice agents on Ultravox: the platform continues. The integration with Inworld’s TTS stack will land over the next few weeks, and the open-weight model positioning (vs the proprietary speech-to-speech stack) is the question worth watching.
ElevenLabs Eleven v4 + Turbo: sub-150ms latency
The third Sep 30 voice-AI move was ElevenLabs shipping Eleven v4 with a Turbo variant that hits sub-150ms latency. The model is positioned as a shift from “reading” to “acting” — i.e., voice agents that participate in real-time conversation rather than narrating pre-scripted content.
TechTimes’ arena benchmark coverage flagged ElevenLabs ahead of Cartesia Sonic 3.6 (scored 1276) with a gap outside the overlapping confidence intervals. The benchmark numbers themselves are secondary; the latency number is the meaningful one. Sub-150ms is the threshold where voice agents stop feeling like a transcription pipeline and start feeling like a phone call. Combined with the dev-tooling story ElevenLabs has been building (Projects, Dubbing Studio, Voice Design), Eleven v4 + Turbo is the production-grade answer to “can voice agents replace call center workflows.”
What the cluster signals
Three voice-AI moves in 24 hours, with no obvious shared trigger, point to three things:
-
Voice agents are now a venture-scale category, not a research frontier. ElevenLabs at $22B, Ultravox acquired by a speech-stack incumbent, sub-150ms production latency — the technical and capital signs both say voice agents are a 2026 product category, not a 2027 frontier. Builders targeting this space should plan on production-grade defaults, not proofs-of-concept.
-
The vertical integration play is on. Inworld’s acquisition of Ultravox is the second voice-AI M&A move in the past quarter. Expect more consolidation as speech-stack companies acquire developer platforms (or vice versa). The customer surface — voice-agent deployments in call centers, IVR, customer support — is where the durable revenue lives.
-
Sub-150ms is the new bar. If your voice agent latency is above 200ms, you’re shipping a prototype. ElevenLabs Turbo is the public benchmark for what “production” means in 2026. Watch for Cartesia, Inworld, and the open-weight speech-to-speech community to respond over the next 90 days.
Sources
- TechCrunch — AI voice startup ElevenLabs doubles valuation to $22B — $22B valuation, $300M tender, Wellington + T. Rowe Price co-leads, $11B Feb 2026 baseline
- Beri.net — Inworld Buys Ultravox and Upgrades Only Its Own Voices — Sep 30 acquisition announcement, Inworld’s built-in voices move to Realtime TTS-2
- Runtimewire — Inworld acquires Ultravox to add a voice-agent platform to its speech stack — deal terms undisclosed, vertical integration analysis
- TechTimes — ElevenLabs Eleven v4 Shifts Voice AI from Reading to Acting, Turbo Hits Sub-150ms Latency — sub-150ms Turbo latency, Cartesia Sonic 3.6 benchmark comparison
- Cryptointegrat — AI News October 1, 2026 — Inworld + Ultravox acquisition confirmation
- ActivePieces — Best Vapi Alternatives — Ultravox open-weight speech-to-speech positioning