TL;DR. Voice launches landed on five straight days. Apple's Siri AI, Google's Gemini 3.8 Live, StepFun's StepAudio 3, DeepL's voice-preserving translation, Speechmatics' agent recogniser, Deepgram's India endpoint, SpaceXAI's Grok Voice Transcribe 2.0 and Alibaba's Qwen3.8-Omni-Flash. On September 16, California made synthetic performer disclosure law.
New models and launches
Apple shipped Siri AI with iOS 27 and macOS 27 (September 14). The assistant gets personal context, onscreen awareness and systemwide app actions, in English beta first, with French, Japanese, Korean, Portuguese and Spanish next month. It needs Apple Intelligence hardware (iPhone 16 models, iPhone 15 Pro, M1 Macs and later), users opt in and may sit on a waitlist, and it is not available in the EU or China at launch. (Apple Newsroom, CNBC, TechCrunch)
Google shipped Gemini 3.8 Live and 3.8 Live Extended Thinking (September 15, post updated September 17). Google says 97 languages with mid-conversation switching, a Speech to Speech Quality Index of 82.6 and 97.7% on Big Bench Audio reasoning for the Extended Thinking model, available in the Gemini API and AI Studio, with 3.8 Live in Search Live. The benchmark numbers are Google's own. (blog.google, Forbes, TechTarget)
SpaceXAI released Grok Voice Transcribe 2.0 (September 18). The company, renamed from xAI in July after SpaceX acquired it, reports word error rate falling from 20.6% to 6.8% on short-phrase testing across 19 languages, at $0.10 per hour of audio batch and $0.20 per hour streaming, with diarization, timestamps and key term biasing included. (xAI, MarkTechPost)
Speechmatics launched Agent STT, a recogniser built for voice agents (September 16). Powered by its Linden 1 model, Speechmatics reports mean finalisation under 350 milliseconds (about 440 at the 95th percentile), 55 or more languages, up to 1,000 custom vocabulary terms, and live diarization. Launch price is $0.30 per hour, falling to $0.16 at volume. The argument behind it is the right one: a low overall word error rate still lets a system get the account number wrong, and the account number is the call. (Speechmatics)
Alibaba released Qwen3.8-Omni-Flash (September 18). A native omni-modal model with a 1 million token context window, priced at $0.15 per million input tokens internationally. Alibaba says audio input costs 98% less per hour than its own Qwen3.5-Omni-Plus, which is a comparison against its predecessor, not against the market. (MarkTechPost, Neowin, BigGo Finance)
StepFun released StepAudio 3, a five model audio family (September 15). Realtime conversation, speech recognition, speech generation, audio generation and music, with the realtime model built for full duplex dialogue, interruptions and tool use. StepFun reports 98.9% on Artificial Analysis Conversational Dynamics, 99.7% on Speech Reasoning and 1.7% word error rate on its recognition model. Those are StepFun's own posted results. (StepFun, StepFun on X)
DeepL Voice began preserving the speaker's voice across languages (September 15). Translated speech now carries the original tone, rhythm and pacing rather than routing every speaker through one synthetic voice, in 12 languages to start, with voice to voice meeting translation generally available and a desktop app spanning Zoom, Teams and Meet. DeepL reports a 4% error rate against a 17% average for the three meeting platforms, a company figure. (DeepL)
Deepgram opened an India endpoint, then shipped a pharmacy model (September 15 and September 17). The India regional endpoint reached general availability for speech to text, text to speech, voice agent and text intelligence workloads, and two days later Nova-3 Pharma arrived, tuned for pharmaceutical vocabulary and drug name recognition. Both are dated entries on Deepgram's public changelog. Read together: where the data sits and which words cannot be wrong are now product decisions. (Deepgram changelog, September 15, Deepgram changelog, September 17)
Voicemod announced the Key Pocket, a hardware voice changer (September 15). Real-time voice changing and soundboards on phones and consoles without a PC, $99.90 for the Founders Edition, shipping late October. Small, physical, and squarely on our early-and-scrappy beat. (Yahoo Finance, The Next Web, audioXpress)
The money
Nuance Labs raised a $50 million Series A (September 14). The Seattle lab, founded in 2025 by former Apple researchers, is building one full duplex model for conversation and expression instead of stitching transcription, language, speech and animation together. Lightspeed led, with NVIDIA and Define Ventures new, Accel and South Park Commons returning. This is Nuance Labs, not the Nuance that Microsoft owns. (Nuance Labs, GeekWire, SiliconANGLE)
Treble raised $18 million in a Series A-2 (September 17). The Icelandic company simulates acoustics and generates synthetic audio data for testing voice-enabled products, and names Amazon and Logitech as users on its own announcement. Paladin Capital Group led, with KOMPAS VC, Frumtak Ventures and the EIC Fund participating, taking total funding to 36 million euros. (GlobeNewswire, Tech.eu, audioXpress)
Law, safety, and consent
California made synthetic performer disclosure law (September 16). Governor Newsom signed SB 1050, the Advertisement Integrity Act, at SAG-AFTRA headquarters. Ads that use AI generated performers must disclose it, audio included, and the rule takes effect on January 1, 2027. (Office of the Governor, Bloomberg Government, MediaPost)
Title firms reported seller impersonation attempts at more than double the 2024 rate (September 14). The American Land Title Association's 2026 Critical Issues Study found 59% of firms saw at least one attempt in the prior year, up from 28% in 2024, and 45% saw one in the prior month, up from 19%. What doubled is reported attempts, not confirmed losses. 87% of firms rated spoofed contact information as at least somewhat common. (HousingWire, Scotsman Guide)
Two Michigan homebuyers lost $66,000 days before closing (September 15). The email looked right, the caller ID matched, and the voice sounded like the mortgage professional they had worked with for weeks. The FBI puts real estate scam losses above $275 million last year, up 58.5% on 2024. (CNN)
Pindrop launched BotStopper, a standalone AI voice agent detector (September 16). The company says 14 of the 20 largest US banks use its technology, and that one Fortune 500 healthcare deployment cut bot activity by 94.3% in under four months after identifying more than 30,000 bot calls in a year. Those figures are Pindrop's. (GlobeNewswire, Pulse 2.0)
Mavenir launched NetAIShield for operator-side fraud (September 17). Fraud protection across voice, messaging and data, aimed at deepfake audio, SIM swap and AI social engineering. Mavenir says its security portfolio runs in more than 70 Tier 1 operator networks. (GlobeNewswire)
In the wild: products, enterprise, and culture
Snap put Specs to work, with Salesforce, AWS and Nvidia (September 16). Agentforce comes to the glasses, AWS connects Amazon Quick so workers can ask for inventory levels by voice, and Nvidia contributes agents from its XR stack. The glasses are $2,195 with a $200 deposit, shipping later this autumn. Several outlets rounded the price to $2,000. (CNBC, Benzinga)
Amazon launched Alexa+ in India (September 16). Early Access for everyone in the country, in English, Hindi and Hinglish, with mid-sentence switching, multi-step tasks and smart home control. It will be free for Prime members after early access and 2,000 rupees a month otherwise. (TechCrunch, Konsulteer)
Oracle Health opened its Clinical AI Agent to inpatient nurses in the US (September 14). Voice driven chart navigation, acute nursing summaries and voice entered discrete charting inside the Foundation EHR. Oracle says the physician version has saved more than 400,000 hours since launch, a company figure. BayCare Health System is an early adopter. (PR Newswire, Fierce Healthcare, Healthcare Dive)
SquadStack moved its voice AI pricing from minutes to outcomes (September 14). A small fixed fee plus a success fee tied to an agreed business result, for example a share of disbursed loan value or a fee per card issued. It is live for banking and financial services sales first. This is the pricing model we have argued for: if you charge for outcomes, spraying calls stops paying. (CIOL, SquadStack)
TiVo introduced Agent TiVo (September 14). Voice or text search for streaming content by mood, genre, actor, sports team or quote, with follow-up questions, routed across models in a hybrid architecture to hold down cost and latency. (TV Tech, Broadband TV News)
DataSync Africa launched AGROSYNC voice advisories in Yoruba (September 15). Nigerian farmers can ask crop questions by voice and get advisories back in Yoruba, with Hausa and Igbo planned. Voice is the right interface when typing is the barrier, and this is the kind of build we want more of. (TechCabal, Business Tech Africa)
Acapela introduced Growing Voices, a synthetic voice that ages with its user (September 14). Built for children who use augmentative and alternative communication, the voice moves through five stages from childhood into adulthood while holding onto the same vocal identity, so a user does not have to swap voices as they grow. The first two are Rosie in UK English and Josh in US English. (Acapela Group)
Wonders.ai opened beta applications for darlin at Tokyo Game Show (September 17). A personal AI companion app launched with voice actor Miki Kohinata, more than ten VTubers and creators, and a Kotobukiya figure collaboration. Strange, specific, and a genuinely different go to market for a voice product. (PR Times, Inside Games)
Also shipped this week, dated by press release: Marvin launched Live Intercept for voice to voice moderated research (Business Wire, September 16), Betterbot added a native CRM alongside its voice AI for after-hours leasing (PR Newswire, September 17), Vimia Tech launched Vosko AI for video localization and dubbing (Send2Press, September 17), FineVoice launched the Aunio audio production agent (Send2Press, September 16), and Guava and CarrierX partnered on 16 kHz HD voice for AI agents (Los Angeles Times, September 14). These are vendor announcements on the wire with no independent reporting behind them.
Papers this week
Benign audio processing pushes deepfake detectors toward "synthetic" (preprint posted September 17). Running genuine human recordings through neural codec analysis and synthesis moved 9 of 13 frozen detectors toward a synthetic verdict, and a 1984 phase retrieval algorithm moved 12 of 13. The sample is small, 47 utterances from 3 speakers, and the preprint is not peer reviewed, but the direction matters: a detector that flags processed human speech is a detector that will accuse real people. (Zenodo)
CrystalASR decomposes recognition into modular layers (preprint posted September 17). A 3.3 million parameter phoneme CTC head, a rule based word decoder and an optional language model reach 17.44% word error rate on LibriSpeech dev-clean, with 21 times fewer trainable parameters and 14 times faster inference than the author's end to end baseline. That error rate is far above modern systems; the claim is efficiency, not accuracy. (Zenodo)
VISH-GUARD screens calls for vishing across languages (published September 18). A multi-agent, LLM-powered framework that combines acoustic features, semantic intent and emotion detection to flag voice phishing in real time. (Scientific Reports)
A skin-conformal inertial sensor reads silent speech (published September 18). The interface sits on a soft substrate and suppresses motion artifacts, recognising speech from lip articulation with no sound produced. Assistive first, but the surveillance question follows it. (ACS Sensors)
DOTA-ME-CS gives code-switching research a dataset (published September 18). 18.54 hours of Mandarin-English code-switching audio, extended with AI timbre synthesis and speed variation. (Journal of Ambient Intelligence and Humanized Computing)
The MultiSOCIAL Toolbox opens up multimodal interaction data (published September 18). Open source extraction of time series from video, including speech transcripts via automatic recognition. (Behavior Research Methods)
A grapheme recall diagnostic exposes Turkic ASR failures (preprint posted September 14). Multilingual recognisers transcribe low-resource Turkic languages in the orthography of better-resourced neighbours, which a word error rate alone hides. (Zenodo)
A conference presentation makes the case on underrepresented languages (posted September 15). Language documentation and preparation as the precondition for digital inclusion. This is a presentation deposit, not a peer reviewed paper. (Zenodo)