TL;DR. Automation Anywhere agreed to buy Boost.ai from Nordic Capital. Amazon made Nova 2.5 Sonic generally available, HeyGen launched its own voice model and Deepgram took Nova-3 Medical to 10 languages. ElevenLabs signed a voice AI agreement with Karnataka and said it will invest hundreds of millions of dollars in India. Giorgia Meloni filed a four-second clip of her voice as an EU trademark, and the FCC is weighing AI-voice political robocalls.


New models and launches

Amazon made Nova 2.5 Sonic generally available (October 5). The speech-to-speech model in Bedrock adds better reasoning, instruction following and tool calling, asynchronous tool calls and a 256K context window. It speaks seven languages, runs in four AWS regions (N. Virginia, Oregon, Stockholm, Tokyo) and costs the same as Nova 2 Sonic.

Source: AWS

 

HeyGen launched HeyGen Voice, its first in-house TTS model (October 9). It is free in the HeyGen platform and API; a Professional Voice Clone add-on costs $99 a month and trains on 30 minutes to 3 hours of speech. It debuted first on Artificial Analysis' Controlled Voice (cloned voice) leaderboard at Elo 1201, ahead of Qwen-Audio-3.1-TTS-Plus (1187) and Eleven v4 Turbo (1168).

Sources: PR Newswire, Artificial Analysis

 

Deepgram took Nova-3 Medical to 10 languages (October 6 to 7). The model now covers Dutch, English, French, German, Hindi, Italian, Japanese, Portuguese, Russian and Spanish. Against Deepgram's general multilingual model, batch WER drops from 6.62% to 4.96% and entity error rate from 14.46% to 9.11%. In Deepgram's comparison, ElevenLabs Scribe v2 Medical scored 5.49% and Microsoft MAI 1.5 scored 7.80%. An upgraded English Nova-3 Medical shipped the same week.

Source: Deepgram

 

TII released Falcon-ASR for Emirati Arabic (October 6). The 1.6 billion parameter speech recognition model transcribes Emirati and Modern Standard Arabic, English, French, Spanish and Portuguese with word-level timestamps. TII says it beat a 30 billion parameter multimodal model on Emirati speech. It shipped with Falcon-Emirati, a 7B language model, and Falcon-OCR-Arabic.

Sources: TII, Semafor

 

Sierra launched fleming-1, a model that detects AI callers (October 8). It scores caller audio in real time for signs the caller is an AI agent, tuned so real people are not flagged, and works with any voice agent built on Sierra. Sierra published no accuracy figures. At its Summit on October 6, Sierra also announced Tandem Voice, which pairs a conversation layer with a reasoning layer, the Curie orchestration model and Persona Studio for custom agent voices.

Sources: Sierra, Sierra Summit recap

 

Hume launched Expression APIs (October 6). The audio API returns 414 vocal expression tags and 190 voice descriptors; the video API covers 48 emotional expression categories and 27 facial descriptions, in 50+ languages. It comes four days after Hume said its TTS and EVI APIs close on November 13.

Source: Hume

 

Sesame gave its voice agents their own computers (October 6). A new preview in the iOS and Android apps adds a new voice model. Each agent can now run tasks on its own computer, orchestrate coding agents such as Devin, Claude Code and Codex, connect to Gmail, Google Calendar and Google Drive, and use custom MCPs. Sesame's glasses are planned for 2027, made in Japan.

Source: Sesame

 

Amazon launched three Alexa tablets with Alexa+ built in (October 8). Alexa Tablet 12 Pro ($499.99), Alexa Tablet 11 ($329.99) and Alexa Tablet 8 ($229.99) run Android with Google Play. Sales start October 14 in the US, Canada and Mexico and October 19 in the UK, Germany, France, Italy and Spain.

Source: Amazon, TechCrunch

 

Aqua Voice brought its dictation app to Android (October 6). It runs on the company's Avalon 1.5 model, which Aqua says scores 5.55% WER on the Open ASR Leaderboard, and dictates in 49 languages.

Source: Business Wire via Yahoo Finance

 

Natura announced Interface, a $99 ring for talking to AI agents (October 8). Press a button, speak, and the request goes to an agent such as Meta's Muse, Claude or ChatGPT. After a free period it costs $9 a month. Preorders open next month, with shipping planned for December or January.

Source: TechCrunch

 

Also shipped this week:

  • Slang AI introduced Superhost, a restaurant phone platform in 97 languages that recognises returning guests (October 5). Source: PR Newswire

  • Invoca launched voice and messaging agents for home services that book jobs in ServiceTitan and Salesforce Field Service (October 6). Source: Invoca

  • CreditNirvana launched CognifAI Assist, multilingual voice guidance for lending journeys (October 6). Source: PR Newswire

  • Rivian brought its AI voice assistant to the R2 in software update 2026.36 (October 7). Source: Electrek

  • Lotte Innovate added LLM voice control to the Lotte Smart Home app for new Lotte Castle and LE EL residences (October 7). Source: Seoul Economic Daily

  • Executable Technologies launched Oria, an assistant that writes in 12 African languages and talks in French and English (October 7). Source: EIN Presswire

  • LigoLab introduced LigoLab Voice for hands-free anatomic pathology work (October 9). Source: GlobeNewswire

 


Infrastructure and tooling

NVIDIA shipped Speech NIM 26.10.0 (October 7). It adds an eight-speaker Sortformer diarization model for Nemotron ASR Streaming and about five times more concurrent streaming sessions than version 1.3.1. NVIDIA lists a known issue: diarization accuracy may drop beyond four speakers.

Source: NVIDIA

 

Telnyx AI Assistants can now run more than one transcription model (October 9). Fallback mode takes up to three backup STT models; booster mode runs a second model alongside the first. Telnyx reports that pairing Deepgram Flux with Nova-3 cut English word errors from 5.7% to 4.6%, and an Arabic and English pair went from 27.7% to 22.4%.

Source: Telnyx

 

SIPPIO gave Microsoft Foundry agents phone numbers (October 8). AI Connect lets agents built in Microsoft Foundry make and receive calls in 80+ countries, with SIPPIO handling local compliance.

Source: SIPPIO via Cloud Communications Alliance

 

Sierra and Meta proposed a Personal Agent Protocol (October 6). The open standard, with Genesys, Instinct, Shopify, Stripe and Walmart among the backers, sets how personal agents deal with businesses across channels, including phone. A v0.1 spec is due later this month.

Source: Sierra

 

OpenAI deprecated its legacy TTS models (October 1, missed last week). tts-1, tts-1-hd and two gpt-4o-mini-tts snapshots leave the API on January 6, 2027. The stated replacement is gpt-realtime-2.1-mini.

Source: OpenAI

 

Also in the changelogs:

  • ElevenLabs made Eleven v4 Turbo the API's default TTS model (October 5). Source: ElevenLabs changelog

  • Vapi added Eleven v4 Turbo and Deepgram nova-3-pharma (October 5). Source: Vapi

  • Deepgram improved Nova-3 for Catalan, French, Indonesian and Vietnamese (October 5). Source: Deepgram changelog

  • LiveKit Agents 1.8.5 added AWS Nova Sonic 2.5 (October 6), and 1.8.6 added Sarvam bulbul:v4-flash and Nabrah Arabic STT (October 9). Sources: GitHub, 1.8.5, GitHub, 1.8.6

  • Sarvam made Bulbul V4 available in its voice agents (October 6). Source: Sarvam

  • Speechmatics agent STT now accepts 8 kHz audio (October 9). Source: Speechmatics changelog

  • Twilio Voice Insights now tags calls that connect in HD audio (October 9). Source: Twilio

 


The money

Automation Anywhere agreed to buy Boost.ai (October 7). The seller is Nordic Capital, which invested in 2021. Terms were not disclosed and the deal is expected to close in Q4 2026. Boost.ai reports 650+ deployments, more than 150 million automated conversations and 36+ languages. It is Automation Anywhere's second AI acquisition in a year, after Aisera.

Sources: PR Newswire, Nordic Capital

 

ElevenLabs said it will invest hundreds of millions of dollars in India (October 6). CEO Mati Staniszewski told Reuters during the company's first India summit in Bengaluru. At the summit, ElevenLabs signed its first agreement with an Indian state, with the Karnataka Innovation and Technology Society. It covers free voice restoration licences for people who lost their speech, Kannada voice access to helplines and citizen services, and a voice AI safety framework. Staniszewski also told CNBC-TV18 he wants the company ready for an IPO "by 2028", with ElevenAgents handling more than 15 million conversations a week.

Sources: The420.in, Reuters, CNBC-TV18 via TradingView

 

People moves:

  • Sarvam AI named former Google Cloud India MD Sashikumar Sreedharan president and chief business officer (October 7). Source: Forbes

  • PolyAI appointed Sunil Shah as CFO (October 8). Source: PR Newswire

 


Law, safety, and consent

Giorgia Meloni filed her voice as an EU sound trademark (October 5). Application 019431219 at the EUIPO is a four-second recording of the Italian prime minister saying "Io sono Giorgia Meloni", in classes 9, 41 and 45. It is under examination. Her office says the aim is protection against AI deepfakes; a mark would cover that recording, not every imitation of her voice.

Sources: PPC Land, Anadolu

 

The FCC is weighing AI-voice political robocalls without consent (comments closed October 5). Club for Growth asked for a waiver so noncommercial political calls to mobile numbers can use "an artificial or prerecorded voice, including an AI-generated voice" without prior consent. Reply comments are due October 19. Commissioner Anna Gomez warned of "a tsunami of robocalls and misinformation".

Sources: FCC, Al Jazeera

 

Royal Mail confirmed its Christmas ad used an AI voice (October 9). Lucy Wynne says the voice in the 2025 ad sounds like her late father, narrator Paul Vaughan. Royal Mail says it "was not designed to replicate any specific individual". The voice came from Czech firm Fameplay using ElevenLabs, and a University of York phonetician found "considerable phonetic similarity". ElevenLabs has added Vaughan to its No Go Voices list.

Sources: Voice Over Herald, Resultsense

 

Mexico arrested five people over AI voice cloning extortion (October 8). The Security Ministry says former Michoacan state treasurer Luis Miranda Contreras, his wife, children and a bodyguard used AI since 2021 to clone the voices of Supreme Court justices, governors, mayors and presidential staff. No total amount was disclosed.

Sources: Expansion Politica, AFP via Arab News

 

ICE asked vendors about voice profiling detainee calls (week of October 5). A request for information describes a system that would build voice profiles from recorded detainee calls, flag calls where the voice does not match the account holder and link calls to location and case records. It is an RFI, not a contract.

Source: Project Salt Box

 

Slovenia's data regulator set rules for AI voice agents in dental offices (October 5). The Information Commissioner's non-binding opinion says practices need a lawful basis for health data, must limit what they collect, tell callers who controls the data and should run a data protection impact assessment.

Source: DataGuidance

 

A US charity plans to monitor Gaza classrooms with AI speech analysis (October 9). Gaza Children Village plans to record conversations in dozens of classrooms and flag "antisemitic", "inciting" or "hateful" speech, keeping recordings for 72 hours. AP reports it is unclear when monitoring will start.

Source: AP via ClickOnDetroit

 

Pegasystems faces a class action over Pega Voice AI (October 2, missed last week). Four plaintiffs in Massachusetts federal court allege the software streamed and transcribed U.S. Bank customer calls to Pega's cloud without consent. These are allegations.

Source: Claim Depot

 

Anthropic started asking Claude users for voice data (October 4, missed last week). A new opt-in asks to use voice chat recordings to improve models. The setting is off by default and separate from the text chat training toggle.

Source: BleepingComputer

 


In the wild: products, enterprise, and culture

Pindrop measured Meta Muse calls reaching contact centers (October 9). Pindrop says it detected Muse agent calls from September 9, a week before Meta announced the beta. Weekly volume grew eightfold by week four, and 64% of the calls reached a human agent. Financial institutions received the most. Pindrop sells detection, and this is its own data.

Source: Pindrop

 

Half of listeners are open to AI-narrated audiobooks, says an Audible survey (October 9). NielsenIQ surveyed 18,000 listeners in 11 countries for Audible: 50% are open to AI narration, from 42% in Germany to 66% in Japan. Audible announced it at the Frankfurt Book Fair.

Source: Voice Over Herald

 

Naver Cloud detected dementia from speech with a 90.14% F1 score (October 1, missed last week). In two Interspeech 2026 papers, the score rose from 81.48% using the transcript alone to 90.14% once narrative topics, speech-flow statistics such as pauses and speech rate, and pronunciation were added. Naver Cloud already runs CLOVA CareCall, an AI service that phones older adults to check in on them.

Source: Naver Cloud

 

General Dynamics will put voice control in Abrams, Stryker and XM30 vehicles (October 7). General Dynamics Land Systems signed a teaming agreement with Primordial Labs to integrate Anura, its natural language command software, which runs at the edge with no cloud.

Source: General Dynamics Land Systems

 

Chick-fil-A said no to AI drive-thru ordering (October 4, missed last week). CEO Andrew Cathy told CNBC: "We're not gonna substitute that interaction with technology." The chain uses AI behind the scenes.

Source: CNBC

 

More voice agents went live in public services and business (October 5 to 9).

  • Seoul's Mapo-gu became the first of the city's 25 districts to answer after-hours resident calls with an AI voice bot. Source: Herald Corp

  • Andhra Pradesh launched a Telugu voice agent for round-the-clock farm advice. Source: New Indian Express

  • Talkaphone turned emergency call stations into information points in 40+ languages. Source: EIN Presswire

  • LAQO, a Croatian insurer, launched a voice agent that calls back people who left a travel insurance purchase unfinished. Source: media-marketing.com

  • Nextech named EliseAI a co-development partner for patient access in specialty practices. Source: Nextech

  • Dyna.Ai and OFFTEC will deploy Arabic voice agents for banks in the Middle East and North Africa. Source: FF News

 


Papers this week

Ambient AI scribes made clinically significant errors in 30% to 68% of surgical transcripts (Journal of Medical Systems, October 9). Eight systems were tested on 100 narrated surgical case reports. Domain WER ran from 3.60% for a GPT-4o Transcribe plus GPT-5 correction pipeline to 24.03% for Amazon Medical Transcribe; errors clustered in medication doses.

Source: Springer

 

Regulators lean on clinician review to govern AI scribes (PLOS Digital Health, October 6). A study of 74 official guidance documents finds clinician review is the central safeguard, without defining the support clinicians need to catch errors.

Source: PLOS

 

Activation steering makes TTS clearer in noise (arXiv, October 6). Steering a pretrained TTS model toward Lombard-style speech, without retraining, cut WER by 7 to 22% at 1 dB SNR while keeping 89 to 95% speaker identity.

Source: arXiv

 

EDICT edits a voice's timbre, then controls delivery segment by segment (arXiv, October 8), with two new benchmarks, TimbreEdit-Bench and IntraTTS-Bench.

Source: arXiv

 

Dialect data cut Lithuanian ASR errors (preprint, October 5). Adding dialect speech to Parakeet-TDT fine-tuning dropped dialect WER from 41.62% to 31.30%. Self-published.

Source: Zenodo

 

Segment-level WER exposes where ASR fails (Computers, October 9). An open-source framework for segment and word-level WER, built around air traffic control speech.

Source: MDPI

 

Automated scoring screens children for language disorder (LSHSS, October 9). On 947 sentence recall recordings, AutoRSR reached .873 sensitivity and .730 specificity with Reverb ASR.

Source: ASHA

 

A small network separates overlapping speech and noise (Scientific Reports, October 8). At 0 dB it reaches 15.60 dB SI-SNR improvement with 3.9 million parameters.

Source: Scientific Reports

 

Hybrid "shallowfake" audio beat forensic checks (Cyber Security journal, October 6). Mixing real and synthetic speech with affordable tools evaded established authentication methods, especially after re-recording.

Source: Henry Stewart

 

Textless Nepali to English speech translation (Journal of Himalaya College of Engineering, October 9). Discrete units and HiFi-GAN reach 42.80 BLEU on FLEURS.

Source: DOI

 

Tamil disfluency correction for transcription (Scientific Reports, October 8). A BERT-BiLSTM tagger reaches 96.57% token accuracy, mostly on synthetic data.

Source: Scientific Reports

 

ASR plus machine translation helped interpreter trainees (Figshare, October 5). In 16 English-Chinese trainees, the combined aid went with higher performance and lower cognitive load.

Source: Figshare

 



Previous edition: Voice AI News, Week 40, September 28 to October 4, 2026