TL;DR. ElevenLabs launched Eleven v4 and closed a $300 million employee tender at a $22 billion valuation. Microsoft, AssemblyAI, Decagon and Inception shipped new speech and voice agent models. Hume is shutting down its TTS and EVI APIs on November 13. Inworld bought Ultravox. Courts in Tokyo and Shanghai ruled that a person's voice is protected against AI cloning.
New models and launches
ElevenLabs launched Eleven v4 and v4 Turbo (September 28). Both models cover more than 90 languages, up from 70. v4 Turbo has a median time to first speech of about 150 ms, and Instant Voice Clones need 10 seconds of audio. Available in ElevenAgents, ElevenCreative and the API. (ElevenLabs, TechCrunch)
Microsoft released MAI-Transcribe-2-Streaming and two MAI-Voice-2.1 models (October 1). The streaming transcription model covers 60 languages with automatic language detection, returns first hypotheses in "just over 100ms", ranks first on Artificial Analysis for final and partial transcripts, and costs $0.54 per audio hour as an introductory rate through the end of 2026. MAI-Voice-2.1 speaks 23 languages and 26 locales at $22 per million characters. MAI-Voice-2.1-Flash costs $15 per million characters and generates 45 seconds of audio in 150 ms end to end. (Microsoft AI, SiliconANGLE, The Decoder)
AssemblyAI shipped Universal-3.6 Pro Realtime (September 29). One model now serves 32 languages with automatic detection, 14 of them new, including Korean, Russian, Persian and Marathi. AssemblyAI reports 5.19% WER on its English voice agent benchmark and a 14.4% entity error rate. Median endpoint latency is unchanged at 537 ms and the price is unchanged at $0.45 per audio hour. (AssemblyAI)
Decagon launched Voice 3 and Chord, its first speech model (October 1). Voice 3 covers 70+ languages and can switch language mid-sentence. Decagon says Chord was trained only on licensed data and consented voice talent, and that about 90% of listeners in its blind tests could not tell Chord from a human speaker. (Decagon, Unite.AI)
Inception released Mercury Voice, a diffusion LLM tuned for voice agents (October 1). Inception reports 320 ms median time to first answer token and 750 ms at p95. That figure covers the LLM stage only, not the full call path. Pricing is $0.40 per million input tokens and $1.50 per million output tokens, halved at launch, and it works with LiveKit, Pipecat, Vapi and Retell. (Inception, Kingy AI)
OpenAI introduced dots, persistent agents you can talk to by voice (September 29, DevDay). Dots run on GPT-6 Astra with their own cloud computer and reach 4,000+ apps through plugins. Users can start a voice call with their dot. Rolling out to Pro and Business Premium first. OpenAI's own pages blocked our fetch; this rests on OpenAI as quoted by 9to5Google. (9to5Google)
Suno opened Speech in beta (October 1). It generates spoken voice and a music bed together as one track, rolling out to all users. (Suno, Martin Cid Magazine)
Tavus previewed Griffin, a real-time conversational video model (October 1). Griffin-Lite generates 720p video in 320 ms chunks at 25 fps, with 0.43 seconds average audio-to-video latency on H100s. It is a research preview for a small group of testers and not available to customers. (Tavus)
Fireflies added free dictation to its desktop apps (September 29). Fireflies Talk types into any text field on Mac and Windows, supports more than 90 languages and stores dictations locally. (Fireflies, TechCrunch)
Audible started letting listeners talk to a character (October 1). In a limited US and UK beta, Character Guide shows who is speaking in real time, starting with Dracula, and Interactive Story opens a live conversation with Renfield. (Audible, TechCrunch)
Sarvam released Saaras V4, speech recognition for all 22 scheduled Indian languages (September 25, missed last week). The model pairs an audio encoder with a 3B hybrid state space model that Sarvam trained from scratch, covers the 22 scheduled languages plus English, and streams first tokens in under 150 ms. (Sarvam, WION)
Also shipped this week: Google moved Vids voiceovers and avatar narration to Gemini 3.8 Flash Lite TTS (Google Workspace Updates, September 29). Mobvoi announced the TicNote Watch, a $249 wrist note-taker that transcribes in 17 languages (PR Newswire, September 30). Mezmo launched Bridge Live, in-person captioning for deaf and hard of hearing professionals (Business Wire via FinancialContent, September 30). SoundHound's generative voice assistant debuted in the Kia Sorento in India (Stock Titan, September 30). LG said it will bring Microsoft's Voice Live speech-to-speech to its ThinQ ON home hub by firmware update this year (BigGo Finance, September 30). Credgenics launched Prix AI, a voice agent platform for collections in 12+ Indian languages (CIOL, October 1).
Infrastructure and tooling
Hume is sunsetting its TTS and EVI APIs (October 2). Access ends November 13, 2026 at 12:01 a.m. EST, and account data is permanently deleted after that date. The APIs stay fully supported until then. Hume gives no reason in the changelog. (Hume changelog)
Tenstorrent and Smallest.ai put voice AI on-premises (October 2). Smallest.ai's Lightning V2 TTS runs on Tenstorrent Galaxy Blackhole servers in the customer's own data center, with usage-based pricing. (Tenstorrent, Unite.AI)
Deepgram shipped three changes in three days (September 30 to October 2). Nova-3 streaming can now swap keyterms mid-stream with a Configure message, no reconnect. Flux TTS gained inline pause markers (500 to 3000 ms) and IPA pronunciation overrides in early access. Self-hosted release 261001 makes Flux TTS watermarking mandatory and removes Whisper support. (Deepgram changelog)
NVIDIA published a recipe for fine-tuning Nemotron ASR on Saudi dialects (September 30). On 133.7 hours of Najdi and Hijazi speech, target-dialect WER fell from 55.05% to 29.96% while English moved from 11.04% to 10.42%, trained in about 4.5 hours on two GPUs. (NVIDIA Developer, Arab News)
Hume tested Google's multi-speaker TTS (September 29). Across five Gemini TTS models and 48 two-speaker scripts, overall scores ran from 3.60 to 4.12 out of 5. Speaker separation scored lowest, and male-male pairs scored 2.87 against 4.65 for mixed-gender pairs. (Hume)
Also in the changelogs: Telnyx added OpenAI's full-duplex GPT-Live to its Voice AI Assistants (Telnyx, September 30) and put an Agent Memory API in beta (Telnyx, October 1). Twilio made Batch Transcription Configuration generally available with Deepgram or a Twilio-managed engine (Twilio, October 1). DeepL moved seven Voice API languages, including Hindi, Tamil and Canadian French, from beta to GA (DeepL, October 1). LiveKit Agents 1.8.4 added Azure Voice Live, a Microsoft AI speech plugin and Eleven v4 support (GitHub, October 1). Bland became an official Twilio technology and go-to-market partner (Bland, September 30). ElevenLabs added transcript editing to speech-to-text and six LLM options to ElevenAgents (ElevenLabs changelog, September 28). Speechmatics reverted earlier latency cuts for 36 realtime languages to restore accuracy (Speechmatics changelog, September 28).
The money
ElevenLabs closed a $300 million employee tender at a $22 billion valuation (September 30). Led by Wellington and T. Rowe Price, with BDT & MSD, EQT, GIC, Goldman Sachs, OTPP and Sapphire Ventures new on the cap table. It is a secondary sale, not new primary capital. The valuation doubled from $11 billion in February. ElevenLabs also opened a Brussels office the same day. (ElevenLabs, TechCrunch, ElevenLabs Belgium)
Inworld acquired Ultravox (announced September 30 and October 1). Terms not disclosed. The Ultravox team joins Inworld and keeps running the platform; built-in Inworld voices on Ultravox move to Realtime TTS-2 at no extra cost. (Inworld, Business Wire via Yahoo Finance)
Modulate raised $25 million (September 28). Led by Future Ventures with Hyperplane and Lakestar, bringing total funding to $60 million. Modulate builds audio-native models for transcription, emotion and deepfake detection. (Modulate, SiliconANGLE, TechCrunch)
Deepslate raised EUR 7.7 million to build European speech-to-speech models (October 1). The Berlin seed round was led by 42CAP with Alstin Capital and SIVentures. (Tech.eu, Tech Funding News)
Presto raised $10 million for drive-thru voice AI (September 28). The money comes "from Remus Capital-affiliated investors and other existing investors." (Business Wire via Yahoo Finance, QSR Web)
Klang raised SEK 15 million and open-sourced a Swedish speech-to-text model (September 23, missed last week). The Helsingborg conversational AI company raised about EUR 1.32 million at a SEK 150 million valuation and released Pianissimo, an open Swedish STT model built on NVIDIA Parakeet. (Klang, Klang Pianissimo)
Law, safety, and consent
A Tokyo court ruled that the human voice is protected by the right of publicity (September 30). Voice actor Kenjiro Tsuda sought removal of 188 AI voice videos posted on TikTok. The Tokyo District Court dismissed the claim because the uploader had deleted the account, but held that a person's voice is protected, the first such recognition in Japan. (The Next Web, Inven Global, AP via WRAL)
A Shanghai appeals court made a voice cloning platform prove where its training data came from (September 29). The Shanghai No. 1 Intermediate People's Court upheld RMB 50,000 in damages for a voice actor. Once she showed a plausible route to her voice and high similarity, the burden shifted to the platform to prove its training data was lawful. (National Law Review, Seoul Economic Daily)
Cerence sued Sony and settled with Microsoft on the same day (October 1). Cerence filed suit in the Eastern District of Texas alleging that Sony's PlayStation and headphones infringe five voice recognition patents. In Delaware, Cerence and Microsoft told the court they had agreed in principle to settle a copyright case over text-to-speech software used after a 2024 licence expiry. Terms are not disclosed and the deal is not final. (Bloomberg Law, Bloomberg Law, Law360)
Resemble AI brought deepfake detection into Microsoft Teams meetings (September 30). Resemble Detect flags audio and video deepfakes in real time and matches face and voice against enrolled people. (Resemble AI)
Rockstar told streamer IlloJuan to drop an AI voice mod (first reported September 29). The mod generated in-game phone calls with cloned voices of real people, including Ibai and Spain's prime minister Pedro Sanchez. Rockstar threatened video removals and copyright strikes, and the mod was removed. (GTA BOOM, Movistar eSports)
In the wild: products, enterprise, and culture
Bengaluru West used AI voice calls to chase property tax (announced September 30). Between March and June, the city corporation placed 4,54,908 calls in Kannada across 64 wards and reached 59,604 owners. 26,962 of them committed to pay Rs 35.11 crore. That is pledged, not collected. The vendor is VocalLabs. (The Bengaluru Live, Asianet Newsable, Deccan Herald)
Airbnb will roll out voice agents this fall (October 1). CEO Brian Chesky said voice agents will help with search and customer service. A stated plan, not a launch. (TechCrunch)
West Virginia is paying for ambient AI scribes in five rural hospitals (September 30). Governor Patrick Morrisey announced $363,000 per hospital for AI ambient clinical listening, part of a $3.2 million Rural Health Transformation package. (WBOY via Yahoo, WDTV)
A 20 second voice sample flagged type 2 diabetes in a study shown at EASD (EASD meeting, Milan, September 28 to October 2). thymia and RMIT University trained on 63,283 samples from 21,129 people. The model caught 82% of cases and wrongly flagged 47% of people without diabetes. A conference presentation, not a peer-reviewed paper. (Medical Xpress)
A "speech clock" estimated biological age from voice (September 30). A Science Advances paper led from Trinity College Dublin's Global Brain Health Institute analysed 2,928 Spanish speakers in five Latin American countries and linked older-sounding speech to markers of brain and biological ageing. (Neuroscience News, Live Science via Yahoo)
AVOXI surveyed 200 contact center and CX leaders (September 30). 55% already use AI voice agents, 17% have fully replaced Tier 1 agents and 60% still run six or more voice providers. AVOXI is a voice provider and ran the survey. (The Fast Mode)
More voice agents went live in service businesses (October 1 to 2). Liberty Home Guard launched Aria for homeowner claims and Ari for service technicians (Business Wire via FinancialContent). Allstar Services extended Revin's voice and SMS agents to 10 more roofing brands after 30% of appointments at Vault Roofing were booked without a human (Yahoo Finance). RealVoice Travel Group introduced TripBeast for vacation rental reservations (Hotel Online). The figures in these releases are the companies' own.
Papers this week
AdaptDuplex lets a full-duplex model decide when to speak (preprint, September 29). Built on Qwen3-Omni, it beats DuplexOmni on 18 of 21 metrics and scores 69.6 on HumDial-FDBench. Self-published, under submission. (Zenodo)
Speculative decoding doubles on-device TTS speed (arXiv, September 29). Real-time synthesis drops from 200 to 88 sequential model calls per second, a 2 to 2.2x speedup for Qwen3-TTS 0.6B on iPhone and Apple Silicon Mac. (arXiv)
RAWD-TTS cuts voice cloning errors (arXiv, September 29). WER falls from 3.18% to 2.42% and speaker similarity rises from 0.733 to 0.748. (arXiv)
Almieyar benchmarks Arabic ASR across 17 dialects (arXiv, September 28). Best system GPT-4o-transcribe at 35.0% WER; Whisper-Large-v3 at 49.5%. (arXiv)
Speech LLM judges reward louder audio (arXiv, September 29). An audit of six speech LLM judges finds they score louder audio higher, a shortcut that matters for anyone using LLM judges to rate TTS. (arXiv)
WenetSpeech-Min releases about 10,000 hours of Minnan speech (arXiv, September 29), with paired Minnan and Mandarin transcripts for ASR and TTS. (arXiv)
Whisper adapted to Macedonian dialects (Applied Sciences, September 29). LoRA on large-v3-turbo cuts WER 57 to 70% against the best zero-shot baseline while updating about 0.7% of parameters. (Applied Sciences)
First neural TTS for Southern Quechua, running on a humanoid robot (Frontiers in Robotics and AI, September 30). Fine-tuned XTTS-v2 on 30.5 hours from 77 speakers; real-time factor 0.961 on a Jetson AGX Orin. (Frontiers)
Deepfake detectors lose accuracy on a new Indian language (preprint, October 2). Single-language baselines lose up to 3.9 EER points on an unseen language; pooling all four languages works best at 0.0325 EER. (Zenodo)
Multi-agent deepfake detection explains its verdicts (Informatics, October 2). A specialist agent targets speaker embeddings; the system detects forged speech, including unseen voice conversion methods, at AUC-ROC 0.79 on MAVOS-DD. (Informatics)
Streaming Kazakh ASR on a 260 hour corpus (October 1). Expanding a 200 hour corpus to 260 hours with voice conversion, a MoChA model reaches 14.8% WER. (NAS RK)
A cheaper pitch detector helps Mandarin cochlear implant users (JASA, October 1). pk-Ctone cuts F0 extraction operations more than 30-fold and improved tone recognition by 10.1 points and word recognition by 19.2 points in a 62 person trial. (JASA)
An alto voice got more drivers to buckle up (Informatics, September 29). In a 12 person preliminary study, seatbelt compliance was 92% with an alto reminder against 58%. (Informatics)
A speech-text model scores kindergarten interactions on embedded hardware (Scientific Reports, October 1). The Mamba-based model runs at 377 ms per utterance and 9.6 W, with an F1 of 0.706. (Scientific Reports)
Speech enhancement lowers Kannada ASR error (International Journal of Speech Technology, September 30). Mean WER fell from 10.15% to 8.73% across 16 acoustic model configurations. (Springer)
MutterMeter detects self-talk through earbuds (ACM IMWUT, September 30). Trained on 34.5 hours from 29 people, it classifies positive and negative self-talk at macro F1 0.815. (ACM)
Previous edition: Voice AI News, Week 39, September 21 to 27, 2026