TL;DR. Google began testing Gemini calls to businesses from your own number, Microsoft opened voice agents in Foundry, and Alibaba, Cartesia, Fish Audio and Google shipped new speech models. Sela raised $21 million for mortgage voice agents. Reuters reported a cloned voice helped move EUR 95 million out of an Intesa unit.
New models and launches
Google let Gemini make phone calls for you (September 24). "Call for Me" is an early preview for paid Gemini subscribers on Pixel 11 phones in the US, in the beta Phone app. Gemini calls a business from your own number to check stock, book a table or reschedule, handles the menus and the hold, and you can read a live transcript and take over. (Google support, TechCrunch, 9to5Google)
Google shipped Gemini 3.8 Flash TTS and Flash-Lite TTS (September 22 in the API changelog, blog post September 23). Flash TTS is pitched at voice design and character work, Flash-Lite at high volume dubbing and voice agents. Google lists 2,000+ production voices, over 100 languages and voice replication from a 30 second sample of a voice you have the rights to use, with built-in consent verification. Voice replication through AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland or India. The #1 placing on Hume AI's Voice Design Benchmark (71.4) is Google's claim. (blog.google, Gemini API changelog, SiliconANGLE)
Google put a face on Gemini 3.8 Live (September 24). Live Avatar adds real-time video avatars with lip sync and turn-taking on top of Gemini 3.8 Live, with SynthID on both audio and video. Gemini Enterprise only, custom avatars by allowlist. (blog.google)
Alibaba shipped Qwen-Audio-3.1 and Qwen3.8-LiveTranslate (September 19 to 23). Qwen-Audio-3.1 is five models: ASR, ASR-Next, TTS, TTS-Next and a realtime model. Qwen says prices drop by about 70% for TTS, about 85% for realtime and up to 95% for ASR; those cuts come from Qwen's own post and we could not read a primary price table. The technical report says task success rose from 78.4% to 82.0% and replies to background speech fell from 73.0% to 13.0% against the previous version. Separately, Qwen3.8-LiveTranslate (September 19) interprets from 60 languages, speaks 29 of them, and cuts average lag from 2.8 to 2.3 seconds by Qwen's measurement. Alibaba Cloud also listed Qwen3.8-Omni-Flash-Realtime on September 21, a realtime endpoint that takes live audio and video and answers in text and audio, with multichannel audio and remote MCP tools. (The Decoder, Alibaba Cloud release list, arXiv tech report, MarkTechPost)
Meta gave Muse a voice mode at Connect (September 23). You design Muse's voice by describing it (faster, slower, an accent). Meta also showed Muse Charm, a pocket device built around a real-time voice model with details due later this year, and Ray-Ban Meta Audio glasses with a 12 hour battery. FDA-cleared hearing enhancement software for mild to moderate hearing loss comes to the US later in 2026 at $149.99. (Meta, Meta Newsroom, TechCrunch)
OpenAI put plugins and GPT-6 models into ChatGPT Voice (September 23). Voice can now run on OpenAI's three GPT-6 models (Astra, Sol and Luna), use plugins such as email, calendar and Slack during a conversation, and start ChatGPT Work tasks on web and mobile, such as drafting a document or summarising Slack messages, then pick them up on desktop. Plus and Pro get Work by voice; Free and Go get plugins and connected apps. OpenAI's own pages blocked our fetch; this rests on OpenAI's announcement as reported. (TechCrunch, 9to5Mac)
Cartesia taught its voices 25 languages (September 23). 50+ library voices now speak up to 25 languages natively, and custom clones from ten seconds of audio can add up to nine accents or languages. (Cartesia)
Fish Audio previewed Drama 3, a TTS model you direct in plain language (September 23). No audio tags or SSML: you describe the delivery, shift voice or style mid-sentence, and regenerate a single word. Gated preview, no pricing. (Fish Audio on X, AlphaSignal)
NKENNEAi released Swahili speech recognition and synthesis (September 25). The first speech models from NKENNE's AI arm, trained on more than 65,000 hours of audio, backed by a $1 million NSF SBIR Phase II award, with a free public demo on September 29. (TechCabal)
Acapela made My-Own-Voice free for Tobii Dynavox users (September 23). People who use AAC devices can bank a personal synthetic voice from a few minutes of recordings, or a donor voice, at no extra cost. It works offline and supports voice banking for children. (Acapela)
SoundHound announced OASYS Edge, agentic voice that runs on the device (September 24). LLM-based voice for cars and smart devices with no cloud connection required, deployment planned for late 2026. (SoundHound)
Navana.ai opened Bodhi TTS (September 24). 50+ voices in 10 Indian languages. (ANI)
Qualcomm announced Snapdragon Sound Elite Gen 2 (September 24). A chip for earbuds that reach cloud agents without a phone. (Gadgets 360)
Qualcomm launched the Snapdragon 8 Elite Gen 6 phone chips (September 22). A new sensing hub runs models of up to 200 million parameters, which Qualcomm says lets a phone run a local scribe that tells speakers apart, and the company says the chip can run a complete voice-in, voice-out agent on the device. (TechCrunch)
Infrastructure and tooling
Microsoft made voice a native agent type in Foundry (September 24). Public preview. Voice agents run on GPT Realtime, Azure Realtime, MAI or your own models, cover 80+ languages and 140+ locales, and deploy to the web, Teams, Teams Phone and Twilio-based inbound and outbound telephony, with voice-specific evaluation and tracing. (Microsoft Azure blog)
NVIDIA open-sourced a diarization model and updated its voice agent blueprint (September 22 and 23). Nemotron 3 Diarization is a 100 million parameter streaming model for up to eight speakers, under the OpenMDW license. NVIDIA reports a 14.72% diarization error rate on VoiceArena's Diarization-Bench, about 22 hours of English conversations. The Voice Agent blueprint v2.2.0 adds an OpenAI Realtime compatible gateway over a cascaded ASR, LLM and TTS pipeline. (Hugging Face model card, GitHub release)
Two vendors published voice benchmarks they built themselves (September 22 and 23). Guava's Daytona model tops Guava's new Voice Index at 58.88 out of 100. The release also reports that evaluators preferred live human agents in 91% of comparisons on response timing and 93% on interruption recovery, across every system tested. Rime released three open TTS benchmarks with 32,450 human judgments across five vendors; Rime leads on support calls, and its data shows ElevenLabs and Deepgram ahead on narration. Both benchmarks are vendor-run; Rime's code is public. (Guava via Business Wire, Rime, Rime on GitHub)
Amazon added bidirectional voice to agent-to-agent collaboration in Connect (September 22). AI agents can now collaborate over the A2A protocol by voice during live customer interactions. (AWS)
AWS End User Messaging added WhatsApp voice calling (September 25). (AWS)
Twilio will fail over Real-Time Transcriptions to a second speech provider (announced September 22, live October 22). If a customer's primary provider does not connect, the session starts on a second provider within the same region. It is on by default for existing accounts and can be switched off per sub-account. (Twilio changelog)
AssemblyAI put Universal-3.5 Pro on OpenRouter (September 22). $0.45 per audio hour. (AssemblyAI)
LiveKit launched Private Links to customer VPCs (September 22). $50 per link plus $0.10 per GB. Not a paid placement. (LiveKit)
Deepgram improved Nova-3 for 11 languages (September 22 and 24). (Deepgram changelog)
Speechmatics sped up batch processing for files up to 20 minutes (September 22). (Speechmatics changelog)
ElevenLabs added parallel tool calls to its agents (September 21). (ElevenLabs changelog)
Applied Brain Research made its on-device ASR and TTS SDK generally available (September 21). (PR Newswire)
Soniox TTS voices arrived on Telnyx (September 25). (Telnyx)
Pipecat 1.12.0 added classifiers for small decisions inside a voice agent (September 26). A classifier answers typed yes/no, choice or score questions about a conversation, such as whether the user's turn is over or whether a greeting is a voicemail, and returns a probability with each answer instead of spending a full LLM turn. JevClassifier runs them on TypeSafe's Jev model, and LLMClassifier uses any Pipecat LLM. Pipecat now grades its own release tests with Jev. The release also adds IPA pronunciation control for Cartesia, ElevenLabs, Inworld and Deepgram Aura-2 voices. (GitHub)
The money
Sela raised $21 million for voice agents in mortgage lending (September 22). Seed and Series A combined, led by Costanoa with Emergence Capital. Sela says its agents help loan officers originate more than $1 billion in mortgages a month, that it passed $10 million in annualised run-rate revenue in 18 months, and that 6 of the 10 largest independent mortgage banks use it. (PR Newswire, FinSMEs)
Heidi raised $340 million for its clinical AI (September 22). A $100 million Series C led by Blackbird plus a $240 million growth investment from General Catalyst. Heidi's products include an ambient listening scribe and voice dictation. (SiliconANGLE, Fierce Healthcare)
DexCare acquired Mila Health (September 21). Mila's agents call, text and chat with patients to get them scheduled. Terms not disclosed. (Business Wire via Morningstar, GeekWire)
Three European rounds, small to smaller (September 23 to 24). Spain's AXIS approved EUR 5 million for Tucuvi, whose voice agent LOLA calls patients for clinical follow-up (Web Capital Riesgo). French white-label voice AI company Reecall raised EUR 3.5 million led by Founders Future and says it has handled over 30 million calls in production (Maddyness). And Dublin's Amethyst Care got EUR 100,000 from Enterprise Ireland's Pre-Seed Start Fund for a voice assistant for older people that handles medication reminders and wellness check-ins (Silicon Republic).
ElevenLabs said it is at $600 million ARR (September 24). CEO Mati Staniszewski gave the figure on stage at Nrth in Toronto, with more than 55% from enterprise. It is a self-reported revenue number, not a new round; TechCrunch reports the company is now valued at $22 billion. He also said there should be disclosure, for now, when customers are talking to an AI agent. (TechCrunch)
InTouchNow raised GBP 2.3 million for AI voice agents in NHS GP practices (September 21). The seed round was led by Ada Ventures with Exceptional Ventures and Kadmos Capital. The agent answers inbound patient calls for booking, triage data collection and admin; the company says it is live in over 150 practices and typically resolves 50 to 70% of calls without a human. (BusinessCloud)
Neosapience debuted on KOSDAQ (September 21). The AI voice company closed its first day at 34,300 won against an IPO price of 10,000 won, up 243%, and says it will put the 20 billion won raised toward North American expansion and its conversational AI business. (BigGo Finance, Chosunbiz)
Law, safety, and consent
A cloned voice helped move EUR 95 million out of Intesa Sanpaolo's Fideuram (reported September 25; the fraud began in February). Per Reuters, fraudsters impersonated Intesa CEO Carlo Messina on WhatsApp, then used an AI clone of a senior law firm partner's voice on a call to confirm transfers to accounts in China and Hong Kong. EUR 53 million has been recovered and EUR 36 million remains untraced. Milan prosecutors have placed a foreign national under investigation. (Reuters via bdnews24, Economic Times)
Consumer advocates asked the FCC not to allow AI-voice political robocalls without consent (September 24). A Club for Growth petition before the FCC would let political campaigns robocall cellphones, AI-generated voices included, without prior consent. NCLC is opposing it ahead of the midterms. (NCLC)
YouTube will add speaking voice detection to likeness protection (September 23). "Later this year," YouTube says, it will combine voice detection with facial detection to catch AI impersonation of creators. (YouTube, The Verge)
Lookout added vishing and voice clone detection to its mobile security platform (September 23). The module analyses call audio and voicemail for cloned voices and flags scam intent from transcripts. Lookout has not published detection rates. (Lookout, Help Net Security)
DaVoice sued Perplexity over wake word technology (filed September 24). DaVoice alleges Perplexity used a proposed partnership to obtain its wake word technology and then built its own. These are allegations in a complaint filed in the Northern District of California. (Bloomberg Law)
The FCC published consumer guides to blocking robocalls (September 22). Carrier and handset specific instructions. (FCC)
Claims opened in Apple's $250 million Siri AI settlement (September 21). The settlement resolves claims that Apple failed to deliver the AI-upgraded Siri it previewed for the iPhone 16. US buyers of an iPhone 15 Pro, 15 Pro Max or any iPhone 16 between June 10, 2024 and March 29, 2025 can file until December 21, 2026, for an estimated $25 per device, which can rise to $95 depending on how many people file. Apple denies wrongdoing. (The Verge, WIRED)
A judge paused discovery in the voice data suit against Nvidia (reported September 25). An Illinois federal judge stayed discovery while he rules on Nvidia's motion to dismiss a class action claiming it used journalists' and voice actors' voices to train AI. (Law360)
In the wild: products, enterprise, and culture
Colorado Springs police put an AI agent on the non-emergency line (September 22). S.A.R.A.H., built on Prepared by Axon, now answers non-emergency calls first, which the city says are over half of all calls to its communications center. It translates in real time across dozens of languages, and callers can still ask for a person. (City of Colorado Springs, KRDO)
Liberty Global signed a three year deal with Sierra, and Virgin Media O2 went first with voice (September 23 and 24). The deal covers about 80 million fixed and mobile connections. Virgin Media O2's voice agent takes selected routine broadband fault calls, a small share of total volume, next to a UK team of 500+ human specialists. (GlobeNewswire, Advanced Television)
NatWest will trial voice-to-voice conversations about your spending (September 25). Royal Bank of Scotland customers go first, on NatWest's own small language model. The bank says its Cora fraud triage agent has already handled nearly 19,000 conversations. (NatWest Group, Finextra)
Deutsche Telekom is replacing phone menus with voice agents (September 24). On stage at HumanX Amsterdam, chief AI officer Kartik Sheth said a new use case starts at no more than 20% of calls handled without a human and reaches 40% within a couple of months. These are his figures, not a published report. (The Next Web)
Drive-thru voice AI moved on two fronts (September 21 and 23). McDonald's told investors it will deploy its ArchIQ system at scale and says its Archy order-taker is above 90% accuracy at early test restaurants, a figure that is McDonald's own and not in the release. Presto joined Toast's partner ecosystem, putting its voice ordering in front of restaurant groups on Toast. (McDonald's via PR Newswire, Semafor, Presto via Yahoo Finance, Restaurant Dive)
Vonage brought branded calling into Epic's patient outreach (September 23). Calls from health systems show the organisation's name, logo and reason for calling, and early adopters report answer rates approaching 45%. (Vonage via Yahoo Finance, HIT Consultant)
The AI receptionist kept spreading (September 22 to 24). Numa launched Operator for car dealerships (PR Newswire), Yuma's agent now completes refunds and address changes during live calls (EIN Presswire), and Yellow Pages launched one for Canadian small businesses (Newswire.ca). The accuracy and automation figures in these releases are the vendors' own.
Netflix used an AI recreation of Gene Wilder's voice in "Wonka's Golden Ticket" (September 23). The Wonka-themed competition show, now streaming, uses the recreated voice as its host; on-screen text says Wilder's estate approved it. (USA Today, Variety)
60 organisations set a five-year goal for AI in underrepresented languages (September 21). Signatories including the Gates Foundation, the OpenAI Foundation, Anthropic, Google, Microsoft, NVIDIA, ElevenLabs and Sarvam aim for an estimated 3.4 billion people who speak languages underrepresented in AI models to be able to use AI tools in their own language and voice. (Gates Foundation)
Papers this week
Cloning patients' voices before tongue cancer surgery (published September 21). A feasibility study cloned the voices of 26 major glossectomy patients across four languages with IndicF5 from a 30 to 60 second preoperative recording, with mean speaker similarity of 0.933 (Resemblyzer) and 0.969 (WavLM). (Frontiers in Oncology)
An ambient scribe tested across 77 languages (preprint posted September 23). In end-to-end tests from speech to clinical note, serious note errors rose from 1% at the simplest complexity level to 24% at the hardest, and 87% of serious errors came from speech recognition. The authors work at Heidi Health, the scribe's maker, and the preprint is not peer reviewed. (medRxiv)
BanglaKontho adds a long-form Bangla TTS corpus (preprint posted September 24). 20 hours and 7,050 utterances from one audiobook narrator. The baseline reaches 9.5% WER and 4.46 MOS, against 16.0% and 3.16 for the same model trained on IndicTTS-Bn. (arXiv)
Inaudible audio can hijack full duplex voice agents (preprint posted September 23). Perturbations hidden under the psychoacoustic masking threshold hijacked an undefended Moshi-style agent in up to 91.7% of white-box trials. The proposed defence cut that to 8.3% at no inference cost. (arXiv)
Deaf and hard of hearing people judge deepfakes differently (preprint posted September 23). In an 80 person study, deaf and hard of hearing participants were 76.4% accurate against 88.0% for hearing participants, mostly by flagging real clips as fake. On audio-only fakes, d/Deaf participants scored 41.2%. (arXiv)
SPADE tests detection of partly edited speech in 12 languages (preprint posted September 21). Localisation models almost always generalise poorly to editing systems they did not see in training. (arXiv)
VietPrism releases 993.4 hours of Vietnamese speech plus 3,100+ hours of matched fakes (preprint posted September 24). Five dialect groups, heavy Vietnamese-English code-switching. A multilingual detector's error rate rose from 16.3% to 33.6% as fake speakers sounded more like the real ones. (arXiv)
Speech recognition for adolescent health in Twi, Dagbani and Ewe (preprint posted September 24). Fine-tuning cut Ewe word error rate from 109.3% to 64.8%. The authors conclude the constraint is validated in-domain data, not the model. (arXiv)
A synthetic conversation harness teaches turn-taking (preprint posted September 23). After fine-tuning on generated two-channel dialogue covering 42 turn-taking phenomena, Moshi took 0.85 of reference turns, up from 0.44. (arXiv)
BanglaTurn detects end of turn in Bangla speech (preprint posted September 24). 84.33% accuracy against 69.28% for the Smart-Turn v3 baseline, at 165 to 191 milliseconds on CPU, with a higher false positive rate as the trade. (arXiv)
YODAS v3 opens 1.1 million hours of multilingual speech (preprint posted September 24). 48 kHz audio in 147 languages under CC BY 3.0. (arXiv)
Previous edition: Voice AI News, Week 38, September 14 to 20, 2026