TL;DR. The week's story is plumbing. LiveKit shipped Connectors, bridging voice agents into Twilio and WhatsApp with no SIP trunk. Google put Gemini Live voice into Gmail, Docs and Keep. Deepgram widened Nova-3's languages, Gupshup went self-serve, and Microsoft repriced transcription to $0.10 per hour of audio. The money was quiet: SoundHound closed its $304M LivePerson deal. Our take: the calling layer is being wired into the platforms people already use, and that matters more than another model demo.

New models and launches

Microsoft released MAI-Transcribe-2 at $0.10 per hour of audio (September 3, 2026). Microsoft says the price is a limited-time offer running through December 31, 2026. The story here is the repricing of the transcription layer, and what the announcement leaves open on enterprise data handling and residency. (microsoft.aiAzure AI Foundry blogVentureBeat)

Google put Gemini Live voice features into Gmail, Docs and Keep (September 3, 2026). Voice-to-document drafting and inbox querying, on paid Workspace tiers. (TechCrunchTech TimesThurrott)

Phonely launched Alma, an LLM built for voice agents (September 1, 2026). Latency and process-following are the claims, and they are the company's own. (No Jitter)

Bodhan AI released four open-weight ASR and TTS models for Indian languages (September 4, 2026). IIT Madras backed. Open weights is the part worth backing. (ET Enterprise AIThe Tribune)

Infrastructure and tooling

LiveKit launched Connectors, bridging voice agents into Twilio Programmable Voice and WhatsApp Business Calling (September 1, 2026). In LiveKit's words, it is "a direct bridge between LiveKit rooms and the platforms people already call", no SIP trunk, no new infrastructure. This is the infrastructure item of the week, and it sits squarely on the outbound-calling beat we care about. (LiveKit)

Deepgram added Kazakh to Nova-3 and improved seven more language models (September 3, 2026). Kazakh ships in both batch and streaming, with upgraded monolingual models for Estonian, Hebrew, Latvian, Lithuanian, Macedonian, Malay and Polish. (Deepgram changelog)

Gupshup launched a self-serve voice AI platform (September 3, 2026). Build, test and deploy voice agents on phone channels without a sales call. (PR NewswireMacau Business)

Havells added multilingual voice control to its connected home line (September 4, 2026). Handling code-mixed and colloquial speech is the interesting part, and it is a vendor-blog claim. (ElevenLabs)

Six VoIP and UCaaS providers rated on in-call intelligence and voice automation (September 4, 2026). Buyer-side reading for anyone picking a stack. (Spiceworks)

The money

SoundHound closed its $304M acquisition of LivePerson and named John Collins CFO (September 4, 2026). The biggest business item of the week. (Business InsiderCMSWireYahoo Finance)

Creoir secured seed investment from Gungnir Capital for defence voice AI (September 3, 2026). Oulu, Finland. The amount is not disclosed on the company's own announcement, so we print no figure. (CreoirKonsulteer)

ElevenLabs named Ashley Kramer Chief Revenue Officer (September 2, 2026). A commercial-side hire, not a round, and it reads as a push toward enterprise sales. (ElevenLabs)

Law, safety, and consent

Voice actors are organising against cloning on two continents (August 30, 2026). The conflict inside Hollywood over AI voice replacement, and the UK "Save Our Voices" campaign backed by Nicola Coughlan, Hugh Bonneville and Matt Lucas, where more than 80 actors have asked the government for statutory rights over their own voices. (Gulf TodayEvening StandardStartup Fortune)

A US law firm accused a UK AI telco vendor of a system that did not work as promised (September 4, 2026). What it looks like when voice automation is oversold. (The Register)

In the wild: products, enterprise, and culture

Job applicants are sending voice agents to their own phone interviews (September 2, 2026). Bot screens bot. This is what happens when the hiring side automates first. (WIRED)

Doctors warned that patient experience is an afterthought in AI scribe rollouts (September 4, 2026). The patient side of a story we covered from the clinician side last week. (The Independent)

A pilot tested weekly wellness calls in the cloned voices of family and friends for older adults (August 31, 2026). A small feasibility pilot, published as a preprint, aimed at loneliness in community-dwelling older adults. (JMIR preprint)

SwitchBot launched MindClip, a wearable that transcribes and summarises conversations (September 3, 2026, global launch around IFA). (FreeYork)

Speechify's CEO on why he buys H100s instead of renting (September 5, 2026). Cliff Weitzman on owning compute, and on what he learned from ElevenLabs. (BigGo)

Lelapa AI says text to speech is next, and is looking beyond Africa (September 4, 2026). An interview and a roadmap rather than a shipped product: the CEO says text to speech is done for three languages so far. (iAfrica)

Deepgram published a guide to paralinguistic cues in voice agents (September 4, 2026). Vendor content, and a useful reference on what prosody and pacing can tell an agent. (Deepgram)

Papers this week

  • Cloned voices were up to 17.5% more intelligible than human voices in noise for middle-aged listeners (September 1). (JASA Express Letters)

  • Accent "cleaning" in ASR reproduces social hierarchies, normalising phonetic variation and excluding speakers with marginalised accents (August 31). (IJEMT)

  • A seven-day in-home Amazon Echo study with visually impaired users found voice autonomy was valued despite recognition errors (September 3). (HFES Proceedings)

  • Four speech-to-text systems compared on Basque subtitles, including Whisper and Speechmatics (September 4). (Perspectives)

  • Deep neural directionality for hearing aids that optimises spatial filtering while preserving environmental awareness (September 4). (Audiology Research)

  • KTU-MEDAFE, a Turkish multimodal dataset with synchronised EEG, speech and facial video from 40 participants (September 3). (Sensors)

  • Voice-based Parkinson's detection from speech signals (September 3). (Bioengineering)

  • Absolutist words in spoken language as a depression marker, tested against clinical diagnoses (August 31). (BMC Psychiatry)

  • P3MER combines federated learning and differential privacy for multimodal emotion recognition (August 31). (Scientific Reports)

  • A survey of large pre-trained speech and language models for Alzheimer's detection (September 3). (OSF preprint)

  • Urdu multimodal emotion recognition reports 91.3% accuracy on the Urdu Speech Emotion Corpus, combining audio and text (September 4). (Scientific Reports)