TL;DR. Automation Anywhere agreed to buy Boost.ai from Nordic Capital. Amazon made Nova 2.5 Sonic generally available, HeyGen launched its own voice model and Deepgram took Nova-3 Medical to 10 languages. ElevenLabs signed a voice AI agreement with Karnataka and said it will invest hundreds of millions of dollars in India. Giorgia Meloni filed a four-second clip of her voice as an EU trademark, and the FCC is weighing AI-voice political robocalls.
New models and launches
Amazon made Nova 2.5 Sonic generally available (October 5). The speech-to-speech model in Bedrock adds better reasoning, instruction following and tool calling, asynchronous tool calls and a 256K context window. It speaks seven languages, runs in four AWS regions (N. Virginia, Oregon, Stockholm, Tokyo) and costs the same as Nova 2 Sonic.
Source: AWS
HeyGen launched HeyGen Voice, its first in-house TTS model (October 9). It is free in the HeyGen platform and API; a Professional Voice Clone add-on costs $99 a month and trains on 30 minutes to 3 hours of speech. It debuted first on Artificial Analysis' Controlled Voice (cloned voice) leaderboard at Elo 1201, ahead of Qwen-Audio-3.1-TTS-Plus (1187) and Eleven v4 Turbo (1168).
Sources: PR Newswire, Artificial Analysis
Deepgram took Nova-3 Medical to 10 languages (October 6 to 7). The model now covers Dutch, English, French, German, Hindi, Italian, Japanese, Portuguese, Russian and Spanish. Against Deepgram's general multilingual model, batch WER drops from 6.62% to 4.96% and entity error rate from 14.46% to 9.11%. In Deepgram's comparison, ElevenLabs Scribe v2 Medical scored 5.49% and Microsoft MAI 1.5 scored 7.80%. An upgraded English Nova-3 Medical shipped the same week.
Source: Deepgram
TII released Falcon-ASR for Emirati Arabic (October 6). The 1.6 billion parameter speech recognition model transcribes Emirati and Modern Standard Arabic, English, French, Spanish and Portuguese with word-level timestamps. TII says it beat a 30 billion parameter multimodal model on Emirati speech. It shipped with Falcon-Emirati, a 7B language model, and Falcon-OCR-Arabic.
Sierra launched fleming-1, a model that detects AI callers (October 8). It scores caller audio in real time for signs the caller is an AI agent, tuned so real people are not flagged, and works with any voice agent built on Sierra. Sierra published no accuracy figures. At its Summit on October 6, Sierra also announced Tandem Voice, which pairs a conversation layer with a reasoning layer, the Curie orchestration model and Persona Studio for custom agent voices.
Sources: Sierra, Sierra Summit recap
Hume launched Expression APIs (October 6). The audio API returns 414 vocal expression tags and 190 voice descriptors; the video API covers 48 emotional expression categories and 27 facial descriptions, in 50+ languages. It comes four days after Hume said its TTS and EVI APIs close on November 13.
Source: Hume
Sesame gave its voice agents their own computers (October 6). A new preview in the iOS and Android apps adds a new voice model. Each agent can now run tasks on its own computer, orchestrate coding agents such as Devin, Claude Code and Codex, connect to Gmail, Google Calendar and Google Drive, and use custom MCPs. Sesame's glasses are planned for 2027, made in Japan.
Source: Sesame
Amazon launched three Alexa tablets with Alexa+ built in (October 8). Alexa Tablet 12 Pro ($499.99), Alexa Tablet 11 ($329.99) and Alexa Tablet 8 ($229.99) run Android with Google Play. Sales start October 14 in the US, Canada and Mexico and October 19 in the UK, Germany, France, Italy and Spain.
Source: Amazon, TechCrunch
Aqua Voice brought its dictation app to Android (October 6). It runs on the company's Avalon 1.5 model, which Aqua says scores 5.55% WER on the Open ASR Leaderboard, and dictates in 49 languages.
Source: Business Wire via Yahoo Finance
Natura announced Interface, a $99 ring for talking to AI agents (October 8). Press a button, speak, and the request goes to an agent such as Meta's Muse, Claude or ChatGPT. After a free period it costs $9 a month. Preorders open next month, with shipping planned for December or January.
Source: TechCrunch
Also shipped this week:
Slang AI introduced Superhost, a restaurant phone platform in 97 languages that recognises returning guests (October 5). Source: PR Newswire
Invoca launched voice and messaging agents for home services that book jobs in ServiceTitan and Salesforce Field Service (October 6). Source: Invoca
CreditNirvana launched CognifAI Assist, multilingual voice guidance for lending journeys (October 6). Source: PR Newswire
Rivian brought its AI voice assistant to the R2 in software update 2026.36 (October 7). Source: Electrek
Lotte Innovate added LLM voice control to the Lotte Smart Home app for new Lotte Castle and LE EL residences (October 7). Source: Seoul Economic Daily
Executable Technologies launched Oria, an assistant that writes in 12 African languages and talks in French and English (October 7). Source: EIN Presswire
LigoLab introduced LigoLab Voice for hands-free anatomic pathology work (October 9). Source: GlobeNewswire
Infrastructure and tooling
NVIDIA shipped Speech NIM 26.10.0 (October 7). It adds an eight-speaker Sortformer diarization model for Nemotron ASR Streaming and about five times more concurrent streaming sessions than version 1.3.1. NVIDIA lists a known issue: diarization accuracy may drop beyond four speakers.
Source: NVIDIA
Telnyx AI Assistants can now run more than one transcription model (October 9). Fallback mode takes up to three backup STT models; booster mode runs a second model alongside the first. Telnyx reports that pairing Deepgram Flux with Nova-3 cut English word errors from 5.7% to 4.6%, and an Arabic and English pair went from 27.7% to 22.4%.
Source: Telnyx
SIPPIO gave Microsoft Foundry agents phone numbers (October 8). AI Connect lets agents built in Microsoft Foundry make and receive calls in 80+ countries, with SIPPIO handling local compliance.
Source: SIPPIO via Cloud Communications Alliance
Sierra and Meta proposed a Personal Agent Protocol (October 6). The open standard, with Genesys, Instinct, Shopify, Stripe and Walmart among the backers, sets how personal agents deal with businesses across channels, including phone. A v0.1 spec is due later this month.
Source: Sierra
OpenAI deprecated its legacy TTS models (October 1, missed last week). tts-1, tts-1-hd and two gpt-4o-mini-tts snapshots leave the API on January 6, 2027. The stated replacement is gpt-realtime-2.1-mini.
Source: OpenAI
Also in the changelogs:
ElevenLabs made Eleven v4 Turbo the API's default TTS model (October 5). Source: ElevenLabs changelog
Vapi added Eleven v4 Turbo and Deepgram nova-3-pharma (October 5). Source: Vapi
Deepgram improved Nova-3 for Catalan, French, Indonesian and Vietnamese (October 5). Source: Deepgram changelog
LiveKit Agents 1.8.5 added AWS Nova Sonic 2.5 (October 6), and 1.8.6 added Sarvam bulbul:v4-flash and Nabrah Arabic STT (October 9). Sources: GitHub, 1.8.5, GitHub, 1.8.6
Sarvam made Bulbul V4 available in its voice agents (October 6). Source: Sarvam
Speechmatics agent STT now accepts 8 kHz audio (October 9). Source: Speechmatics changelog
Twilio Voice Insights now tags calls that connect in HD audio (October 9). Source: Twilio
The money
Automation Anywhere agreed to buy Boost.ai (October 7). The seller is Nordic Capital, which invested in 2021. Terms were not disclosed and the deal is expected to close in Q4 2026. Boost.ai reports 650+ deployments, more than 150 million automated conversations and 36+ languages. It is Automation Anywhere's second AI acquisition in a year, after Aisera.
Sources: PR Newswire, Nordic Capital
ElevenLabs said it will invest hundreds of millions of dollars in India (October 6). CEO Mati Staniszewski told Reuters during the company's first India summit in Bengaluru. At the summit, ElevenLabs signed its first agreement with an Indian state, with the Karnataka Innovation and Technology Society. It covers free voice restoration licences for people who lost their speech, Kannada voice access to helplines and citizen services, and a voice AI safety framework. Staniszewski also told CNBC-TV18 he wants the company ready for an IPO "by 2028", with ElevenAgents handling more than 15 million conversations a week.
Sources: The420.in, Reuters, CNBC-TV18 via TradingView
People moves:
Sarvam AI named former Google Cloud India MD Sashikumar Sreedharan president and chief business officer (October 7). Source: Forbes
PolyAI appointed Sunil Shah as CFO (October 8). Source: PR Newswire
Law, safety, and consent
Giorgia Meloni filed her voice as an EU sound trademark (October 5). Application 019431219 at the EUIPO is a four-second recording of the Italian prime minister saying "Io sono Giorgia Meloni", in classes 9, 41 and 45. It is under examination. Her office says the aim is protection against AI deepfakes; a mark would cover that recording, not every imitation of her voice.
The FCC is weighing AI-voice political robocalls without consent (comments closed October 5). Club for Growth asked for a waiver so noncommercial political calls to mobile numbers can use "an artificial or prerecorded voice, including an AI-generated voice" without prior consent. Reply comments are due October 19. Commissioner Anna Gomez warned of "a tsunami of robocalls and misinformation".
Sources: FCC, Al Jazeera
Royal Mail confirmed its Christmas ad used an AI voice (October 9). Lucy Wynne says the voice in the 2025 ad sounds like her late father, narrator Paul Vaughan. Royal Mail says it "was not designed to replicate any specific individual". The voice came from Czech firm Fameplay using ElevenLabs, and a University of York phonetician found "considerable phonetic similarity". ElevenLabs has added Vaughan to its No Go Voices list.
Sources: Voice Over Herald, Resultsense
Mexico arrested five people over AI voice cloning extortion (October 8). The Security Ministry says former Michoacan state treasurer Luis Miranda Contreras, his wife, children and a bodyguard used AI since 2021 to clone the voices of Supreme Court justices, governors, mayors and presidential staff. No total amount was disclosed.
Sources: Expansion Politica, AFP via Arab News
ICE asked vendors about voice profiling detainee calls (week of October 5). A request for information describes a system that would build voice profiles from recorded detainee calls, flag calls where the voice does not match the account holder and link calls to location and case records. It is an RFI, not a contract.
Source: Project Salt Box
Slovenia's data regulator set rules for AI voice agents in dental offices (October 5). The Information Commissioner's non-binding opinion says practices need a lawful basis for health data, must limit what they collect, tell callers who controls the data and should run a data protection impact assessment.
Source: DataGuidance
A US charity plans to monitor Gaza classrooms with AI speech analysis (October 9). Gaza Children Village plans to record conversations in dozens of classrooms and flag "antisemitic", "inciting" or "hateful" speech, keeping recordings for 72 hours. AP reports it is unclear when monitoring will start.
Source: AP via ClickOnDetroit
Pegasystems faces a class action over Pega Voice AI (October 2, missed last week). Four plaintiffs in Massachusetts federal court allege the software streamed and transcribed U.S. Bank customer calls to Pega's cloud without consent. These are allegations.
Source: Claim Depot
Anthropic started asking Claude users for voice data (October 4, missed last week). A new opt-in asks to use voice chat recordings to improve models. The setting is off by default and separate from the text chat training toggle.
Source: BleepingComputer
In the wild: products, enterprise, and culture
Pindrop measured Meta Muse calls reaching contact centers (October 9). Pindrop says it detected Muse agent calls from September 9, a week before Meta announced the beta. Weekly volume grew eightfold by week four, and 64% of the calls reached a human agent. Financial institutions received the most. Pindrop sells detection, and this is its own data.
Source: Pindrop
Half of listeners are open to AI-narrated audiobooks, says an Audible survey (October 9). NielsenIQ surveyed 18,000 listeners in 11 countries for Audible: 50% are open to AI narration, from 42% in Germany to 66% in Japan. Audible announced it at the Frankfurt Book Fair.
Source: Voice Over Herald
Naver Cloud detected dementia from speech with a 90.14% F1 score (October 1, missed last week). In two Interspeech 2026 papers, the score rose from 81.48% using the transcript alone to 90.14% once narrative topics, speech-flow statistics such as pauses and speech rate, and pronunciation were added. Naver Cloud already runs CLOVA CareCall, an AI service that phones older adults to check in on them.
Source: Naver Cloud
General Dynamics will put voice control in Abrams, Stryker and XM30 vehicles (October 7). General Dynamics Land Systems signed a teaming agreement with Primordial Labs to integrate Anura, its natural language command software, which runs at the edge with no cloud.
Source: General Dynamics Land Systems
Chick-fil-A said no to AI drive-thru ordering (October 4, missed last week). CEO Andrew Cathy told CNBC: "We're not gonna substitute that interaction with technology." The chain uses AI behind the scenes.
Source: CNBC
More voice agents went live in public services and business (October 5 to 9).
Seoul's Mapo-gu became the first of the city's 25 districts to answer after-hours resident calls with an AI voice bot. Source: Herald Corp
Andhra Pradesh launched a Telugu voice agent for round-the-clock farm advice. Source: New Indian Express
Talkaphone turned emergency call stations into information points in 40+ languages. Source: EIN Presswire
LAQO, a Croatian insurer, launched a voice agent that calls back people who left a travel insurance purchase unfinished. Source: media-marketing.com
Nextech named EliseAI a co-development partner for patient access in specialty practices. Source: Nextech
Dyna.Ai and OFFTEC will deploy Arabic voice agents for banks in the Middle East and North Africa. Source: FF News
Papers this week
Ambient AI scribes made clinically significant errors in 30% to 68% of surgical transcripts (Journal of Medical Systems, October 9). Eight systems were tested on 100 narrated surgical case reports. Domain WER ran from 3.60% for a GPT-4o Transcribe plus GPT-5 correction pipeline to 24.03% for Amazon Medical Transcribe; errors clustered in medication doses.
Source: Springer
Regulators lean on clinician review to govern AI scribes (PLOS Digital Health, October 6). A study of 74 official guidance documents finds clinician review is the central safeguard, without defining the support clinicians need to catch errors.
Source: PLOS
Activation steering makes TTS clearer in noise (arXiv, October 6). Steering a pretrained TTS model toward Lombard-style speech, without retraining, cut WER by 7 to 22% at 1 dB SNR while keeping 89 to 95% speaker identity.
Source: arXiv
EDICT edits a voice's timbre, then controls delivery segment by segment (arXiv, October 8), with two new benchmarks, TimbreEdit-Bench and IntraTTS-Bench.
Source: arXiv
Dialect data cut Lithuanian ASR errors (preprint, October 5). Adding dialect speech to Parakeet-TDT fine-tuning dropped dialect WER from 41.62% to 31.30%. Self-published.
Source: Zenodo
Segment-level WER exposes where ASR fails (Computers, October 9). An open-source framework for segment and word-level WER, built around air traffic control speech.
Source: MDPI
Automated scoring screens children for language disorder (LSHSS, October 9). On 947 sentence recall recordings, AutoRSR reached .873 sensitivity and .730 specificity with Reverb ASR.
Source: ASHA
A small network separates overlapping speech and noise (Scientific Reports, October 8). At 0 dB it reaches 15.60 dB SI-SNR improvement with 3.9 million parameters.
Source: Scientific Reports
Hybrid "shallowfake" audio beat forensic checks (Cyber Security journal, October 6). Mixing real and synthetic speech with affordable tools evaded established authentication methods, especially after re-recording.
Source: Henry Stewart
Textless Nepali to English speech translation (Journal of Himalaya College of Engineering, October 9). Discrete units and HiFi-GAN reach 42.80 BLEU on FLEURS.
Source: DOI
Tamil disfluency correction for transcription (Scientific Reports, October 8). A BERT-BiLSTM tagger reaches 96.57% token accuracy, mostly on synthetic data.
Source: Scientific Reports
ASR plus machine translation helped interpreter trainees (Figshare, October 5). In 16 English-Chinese trainees, the combined aid went with higher performance and lower cognitive load.
Source: Figshare
Previous edition: Voice AI News, Week 40, September 28 to October 4, 2026