TL;DR. OpenAI shipped GPT-Live-1, a full-duplex voice model, and Yelp and Hatch are already running it in production. Apple put ambient listening on the wrist. In the same seven days, China's top court wrote the first judicial rules on cloned voices and Anthropic documented a voice-cloning propaganda operation. Our take: capability and consent arrived together, and a new paper puts synthetic voice mostly in lead generation, not impersonation, so the line is still whether the person wanted the call.
New models and launches
OpenAI shipped GPT-Live-1, a full-duplex voice model, in its API (September 10, 2026). It listens and speaks at the same time, so callers can interrupt or change direction mid-sentence, and it launched with 12 built-in voices. (OpenAI, The Register, The New Stack)
Apple put Siri Recap, Live Rewind and Sound Recognition on Watch Series 12 and Ultra 4 (September 9, 2026). Siri Recap summarises recent conversation, Live Rewind gives instant text of the last 15 seconds of speech, and Sound Recognition flags sounds for accessibility. The story is the one TechCrunch named, technology that is always listening becoming normal, not the spec sheet. (TechCrunch, Axios, WIRED)
Apple is adding five languages to Siri in October (announced September 9, 2026). French, Japanese, Korean, Portuguese and Spanish, with stronger on-device models on newer iPhones. Separate rollout from the Watch features, announced the same day. (Engadget, Dataconomy)
Meta launched Muse, its personal AI agent, in the US (September 8, 2026). The underlying Muse Voice Transcribe model, the one that decides every 80 milliseconds and is aimed at always-on glasses, shipped on September 1 and is context here, not this week's news. The custom voices reported around the launch have no ship date. (Meta, TestingCatalog)
Infrastructure and tooling
Telcos moved on voice AI from three directions (September 9 to 10, 2026). Alianza launched its Crux platform to help operators monetise AI voice (September 9), Radisys launched a V.AI ecosystem for the same goal (September 10), and Deutsche Telekom outlined AI voice services for business customers, framed in its own coverage as a plan for coming months rather than a live rollout. Fierce Network's companion piece, arguing telco adoption is still an uphill climb, is the honest counterweight to print alongside the launches. (Fierce Network on Crux, Business Wire on Radisys, Telco Titans on Deutsche Telekom, Fierce Network on adoption)
Alianza acquired Skribby the day after launching Crux (September 10, 2026). (Markets Insider)
Phony.ai launched a provider-neutral AI phone and voice agent platform (September 10, 2026). Businesses pick their own carrier, language model and speech provider instead of a bundled stack. This is wire-distributed company copy with no independent reporting behind it, so read it as a launch, not as validation. (Markets Insider, Macau Business)
Intermedia launched AI Receptionist, a native voice agent in its digital teammates line (September 10, 2026). (GlobeNewswire via Globe and Mail, Telecom Reseller)
CallRail shipped real-time HubSpot scheduling, turning inbound calls into booked appointments(September 9, 2026). Announced at Unbound 2026. (PR Newswire, CustomerThink)
Dialpad partnered with Rime to make its AI phone agents sound more human (September 10, 2026). (Rime, citybiz)
Twilio introduced orchestration tools for AI-powered customer voice conversations (week of September 8, 2026). We could not pin the announcement to a single day, so we print the week. (MarketBeat)
The money
Listen Labs walked away from a signed $125M Series C at a $1.5B valuation while Salesforce held separate talks to acquire it for around $2B (September 9, 2026). Two numbers describing two different things: the term sheet it turned down and the acquisition conversation it is reportedly still having. Neither is a closed deal, and the company did not comment. (TechCrunch)
Universal Music Group and ElevenLabs are launching an AI music creation platform (September 11, 2026). The licensing angle is the part that matters: it turns an AI speech company into a paying customer of the catalogue. (Semafor, Mixmag)
Pocket FM doubled its annualised revenue run rate to $500M (September 10, 2026). Its own text-to-speech powers 93% of the existing catalogue and 99% of new content. A real revenue number attached to synthetic voice. (TechCrunch)
Japan's nocall raised about ¥800M, roughly $5.2M, in a Series A to automate phone work (September 11, 2026). We print the yen figure because that is what the primary announcement states. (PR TIMES, Dealroom)
Law, safety, and consent
China's Supreme People's Court issued its first judicial rules on AI disputes, covering unauthorised voice cloning (September 7, 2026). Civil liability for deepfakes, cloned voices and algorithmic price discrimination, across 24 articles. This is a new guideline, not a repeat of the short-drama cloning story. (State Council Information Office, SCMP, HKFP, Digital Watch)
Anthropic's threat intelligence report documents Claude being used for live impersonation and a voice-cloning propaganda operation (September 10, 2026). We scope this to the voice findings. The report's separate weapons-related findings are outside our beat. (Anthropic, Axios)
Scammers cloned Bangladeshi lawmakers' voices and hacked their phone numbers to solicit urgent money transfers (September 10, 2026). Logged as an incident by the OECD's monitor, not a press release. (OECD AI Incidents Monitor)
Voice-cloning scams kept working on families (reported September 8, 2026). A Crestview, Florida woman nearly lost $25,000 to a jury-duty voice clone, and security researchers separately warned that scammers are cloning children's voices to pressure parents. (WSMV, WEAR-TV)
EA Sports cloned commentator John Buccigross's voice for NHL 27, with his oversight (confirmed by EA on September 11, 2026). The consented version of the same technology behind the three items above. (Gadget Review, IXBT.games)
In the wild: products, enterprise, and culture
Yelp and Hatch put GPT-Live-1 into production for restaurants and service pros (September 9, 2026). More than a million calls in, customers can interrupt the assistant or switch language mid-call. The real-world proof point under the launch above. (Business Wire, PPC Land)
Xactly's Penny coaches managers through difficult conversations by voice (September 12, 2026). Tailored scripts and real-time feedback on tone during role-play, ahead of the performance and pay conversations the tool is built for. (CNBC)
OnlyFans promoters on X are using AI-generated voice notes to pass as human (reported September 2026). Synthetic voice deployed to manufacture trust at scale rather than to defraud outright, which is the quieter half of the spam problem. (Malwarebytes)
Mountain View is piloting real-time AI translation at City Council meetings (reported September 9, 2026). Civic use, where the question is access rather than consent. (Mountain View Voice)
An AI-narrated audiobook by Terence Ang, a stroke survivor living with aphasia, won a 2026 NYC Big Book Award (2026 winners list, Audiobook: Poetry and Memoir). The good-news item of the week. (NYC Big Book Award winners, Des Moines Register)
Papers this week
Nine papers, each verified against arXiv or Crossref.
Synthetic and automated voices in telemarketing are concentrated in lead generation, not impersonation fraud. The paper behind our take this week. (arXiv)
A deep-fake CAPTCHA framework to defend against real-time voice and face cloning in social engineering attacks. (arXiv)
ZipCodec, an ultra-low-frame-rate streaming speech coding model, evaluated on speech recognition and speaker identification. (arXiv)
SphereVAE, balancing acoustic recoverability against downstream sequence modelling in LLM-based speech generation. (arXiv)
Fine-grained error correction improves Korean speech recognition in call-centre conditions. (arXiv)
Edge-deployed voice language model pipelines evaluated for speakers with atypical language patterns, measuring how recognition noise propagates. (Nature Communications)
Treating conversational AI companions as social entities rather than tools risks emotional dependence and psychological harm. (JMIR Mental Health)
An AI-guided health education conversational agent for gastric cancer patients, built on the OpenMEDLab 2.0 foundation model. (JMIR)
AI-generated podcasts still miss the conversational rhythm of human speakers, despite realistic voices. (AI & Society)