• ElevenLabs Seeks Twenty Two Billion Dollar Valuation — The voice AI startup is in early talks for a secondary share sale that would double its valuation to $22 billion. (Tech Times)

  • Google Maps Updates New Zealand Voice — Google Maps rolled out a new English voice with a Kiwi accent, trained by Te Taura Whiri, that accurately pronounces indigenous Maori place names across Android, iOS, and CarPlay. (blog.google)

  • Amazon Upgrades Alexa For Screenless AI Devices — Amazon's devices chief Panos Panay discussed the company's push into screenless, AI-enabled gadgets, confirming its custom AZ3 chips handle wake words locally while more complex processing is routed to the cloud. (cnbc.com)

  • xAI Launches Grok Voice Agent Builder — xAI released its Voice Agent Builder in beta, a no-code platform that allows enterprises to build custom AI voice agents and phone call centers in minutes using cloned human voices. (xAI)

  • Grok Voice Mode Launches On Apple CarPlay — CarPlay users can now access xAI's Grok assistant as an alternative voice option in their vehicles. (Mashable)

  • SpaceXAI Launches Voice APIs On Vercel Gateway — xAI launched its voice APIs, featuring realtime voice, text-to-speech, and speech-to-text models, on the Vercel AI Gateway. (BASENOR)

  • Netflix Recreates Gene Wilder Voice With AI — Netflix partnered with ElevenLabs and the late actor's estate to recreate Gene Wilder's voice for a new Willy Wonka-themed reality competition series, drawing both family support and fan backlash. (reuters.com)

  • Meta Caps Conversational Glasses Feature Usage — Meta introduced a three-hour monthly limit on its Conversation Focus feature for smart glasses, locking extended access behind a $20-a-month Meta One Premium subscription — despite the feature running entirely on-device. (PCMag Middle East)

  • Samsung Interactive Displays Auto Transcribe Classroom Lessons — Samsung Electronics introduced new AI features for its educational interactive displays at ISTE Live 26, allowing teachers to automatically convert spoken lessons into text. (Seoul Economic Daily)

  • FBI Warns Of Rising Voice Cloning Scams — The bureau reported that AI-enhanced fraud, including voice cloning and deepfake impersonations, cost Americans $893 million, as new defenses emerge. (MSN)

  • India Launches Open Source VoicERA Voice Stack — The Ministry of Electronics and Information Technology unveiled an open-source, end-to-end voice AI stack built on the BHASHINI infrastructure. (DD News)

  • Papua New Guinea Introduces Voice Cloning Framework — The government commenced policy and legislative frameworks to protect citizens from AI deepfakes, voice cloning, and digital impersonation. (Post Courier)

  • San Diego Police Deploy AI Call Systems — Two local law enforcement agencies are using AI to transcribe and route non-emergency calls. (FOX 5 San Diego & KUSI News)

  • Trump Shares Deepfake Voice Video — A video posted by Donald Trump features an AI-generated deepfake voice mimicking him as a doctor diagnosing his critics. (forbes.com)

  • SoundHound AI To Acquire LivePerson Platform — SoundHound AI announced plans to acquire customer engagement company LivePerson to drive its enterprise expansion. (Foreign Policy Journal)

  • Deepgram Reaches One Point Three Billion Valuation — The voice AI infrastructure company has raised over $240 million to date as it continues to enable software to communicate with humans. (techcrunch.com)

  • ElevenLabs Launches One Million Voices Initiative — The voice AI company announced a pledge to restore one million voices using artificial intelligence. (Mashable)

  • Retell AI Launches Conductor For Voice Agents — Voice AI startup Retell AI introduced Conductor, a graph-native review interface designed to help enterprises test, evaluate, and improve production voice agents. (Yahoo Finance)

  • Hugging Face Launches Open Speech Pipeline — The company partnered with Cerebras to release an open-source, modular speech-to-speech AI pipeline built on Gemma 4 offering sub-second latency for enterprise deployments. (Let's Data Science)

  • SoundHound AI Automates Restaurant Voice Ordering — SoundHound AI is bringing voice and agentic AI to daily restaurant operations, including drive-thru automation, ordering, and guest services. (Restaurant Technology News)

  • AI Receptionist Startup Pie Secures $19.5M — AI receptionist startup Pie has raised $19.5 million amid shifting state laws regarding automated bot disclosures. (Tech Times)

  • Lucida AI Secures €6.1M Seed Round — London-based speech-to-speech AI startup Lucida AI raised €6.1 million to expand its global communication platform. (BeBeez International)

  • Pocket Raises $11M for AI Note-Taking Device — The startup secured $11 million in funding from investors, including ElevenLabs CEO Mati Staniszewski, to expand its phone-attachable voice recording and transcription device. (techcrunch.com)

  • ZF Launches AI Voice Agent For Workshops — The new tool adds real-time, natural-language phone interaction to ZF's management platform to improve customer intake and scheduling. (ZF Press Center)

  • BMW Financial Services Plans AI Voice Tech — The company plans to roll out customer-facing AI voice technology for its call centers in the fourth quarter. (Auto Finance News)

  • RingCentral Adds Native AI Agents To RingCX — RingCentral integrated native AI agents into its RingCX customer engagement portfolio to automate outreach and manage intelligent handoffs. (MarTech Cube)

  • JCL Credit Leasing Deploys Voice Agents — JCL Credit Leasing commercially rolled out real-time AI voice agents from AI Rudder for customer verification and debt collection across its operations in Malaysia. (PR Newswire Asia)

  • Adira Finance Reduces Costs With AI Rudder — Adira Finance successfully lowered its operational costs by implementing AI Rudder's artificial intelligence solutions. (The Manila Times)

  • RevComm Launches MiiTel For Retail Voice AI — The new service records and analyzes face-to-face conversations at stores, service counters, and field locations. (International Business Times)

  • Exotel Powers India's First Agentic Voice Call — The company successfully initiated the country's first fully agentic AI voice call for enterprise customer experience. (Express Computer)

  • Gnani.ai Scales Up Agentic Voice AI Strategy — Bengaluru-based voice-first agentic AI startup Gnani.ai is expanding its operations after overcoming early market hurdles. (Deccan Herald)

  • iFlytek Launches Offline Voice Translation Device — iFlytek introduced the Translator X3 to provide near-real-time offline speech translation for travelers and teams. (AD HOC NEWS)

  • Shilo Launches Voice AI Sales Coach — Shilo introduced an autonomous voice-to-voice AI sales coach that uses agent call data to provide personalized training. (HousingWire)

  • Grido Launches Voice First Mental Wellbeing Companion — Grido AB launched its voice-first AI mental wellbeing companion on the App Store across the European Union. (Fitt Insider)

  • Nebius Debuts Voice Controlled Echo Assistant — Nebius introduced its Aether 3.6 platform featuring a voice-activated AI assistant named Echo. (AD HOC NEWS)

  • ViiTorVoice Text To Speech Model Goes Open Source — The new non-autoregressive model allows users to edit single words inside finished audio recordings without full regeneration. (Tech Times)

  • Interfaze Open Sources Diffusion ASR Model — The startup released a small speech recognition model capable of transcribing six languages using a parallel denoising decoder. (MarkTechPost)

  • Bosonai Releases Expressive Foundation TTS Model — Bosonai built a 2.3-billion-parameter text-to-speech foundation model called Higgs-tts-2-3b-base to generate natural-sounding speech. (HackerNoon)

  • SoundWise Launches Free AI Transcription Tool — SoundWise announced a free browser-based AI transcription tool that supports local processing and over 98 languages. (MarTech Cube)

  • Vmake Labs Launches AI Video Translator — Vmake Labs introduced an AI video translation workflow featuring dubbing, lip-syncing, and voice matching. (The Manila Times)

  • Ubuntu Develops Speech To Text Feature — Ubuntu is working on an AI-powered speech-to-text transcription project called Myna to allow users to talk instead of type. (OMG! Ubuntu)

  • Reachy Mini Robot Integrates Local Voice AI — Hugging Face's open-source speech-to-speech library now powers fully local conversational AI on the Reachy Mini desktop robot. (Hackaday)

  • Revin Partners With Vertex Service Partners — Revin partnered with Vertex Service Partners to deploy its AI voice and SMS agents across roofing operations. (The Fast Mode)

  • Syntiant Partners With Vibe For Voice — Syntiant and Vibe collaborated on a low-power, always-on physical AI solution to enable voice-driven workspace experiences. (The Manila Times)

  • Iberika Telecom Launches AI Voice Agent — Malaga-based Iberika Telecom launched an AI voice agent designed to reformulate traditional business customer service. (Demócrata)

  • Brothers Build Restaurant Voice AI Startup — Two brothers from North Carolina launched a startup focused on automating restaurant operations using native AI voice technology. (GrepBeat)

  • IntelePeer Wins Voice AI Technology Excellence Award — The company's SmartAgent Appointment Scheduler was recognized for advancing conversational AI in healthcare. (MarTech Cube)

  • AI Voice Giant Faces Severe IPO Hurdles — An unnamed AI voice giant is facing IPO difficulties after distributors accused the company of overpromising and delivering products with quality control issues. (36Kr)

  • AI Voice Assistant Prepares Cardiac Patients — A prospective study evaluated Sofiya, an AI voice assistant that successfully called patients to provide pre-procedural instructions and collect clinical data. (npj Digital Medicine)

  • Cross-Subject Brain-to-Text Decoder Developed — Scientists presented a neural-to-phoneme decoder trained on two large intracortical speech datasets to improve cross-subject generalization for speech brain-computer interfaces. (Journal of Neural Engineering)

  • Evaluating Medical Speech Transcription Quality — A study demonstrated that disagreement among heterogeneous automatic speech recognition systems can serve as a reference-free signal to localize transcription uncertainty. (Frontiers in Artificial Intelligence)

  • Deepfakes Threaten Voice Biometric Authentication Systems — A systematic review analyzes key vulnerabilities in face and voice recognition systems to propose a layered biometric security framework against deepfake attacks. (Merkurius Jurnal Riset Sistem Informasi dan Teknik Informatika)

  • Controllable Emotional Speech Synthesis Framework — Researchers proposed the Dual-Track Residual Framework to add controllable emotional expressions to a frozen text-to-speech backbone while preserving speaker identity. (Applied Sciences)

  • Speech Emotion Recognition Model Proposed — A new speech emotion recognition model integrates the WavLM pre-trained model with a multi-head attention mechanism to capture emotionally salient time segments. (Electronics)

  • Arabic Speech Emotion Recognition Reviewed — A systematic review of Arabic speech emotion recognition research from 2015 to 2024 analyzed 24 datasets, dialectal variations, and classification methods. (PeerJ Computer Science)

  • Dual-Mode Systems Advance Speech Emotion Recognition — A literature review examines recent advancements in emotion recognition systems that combine facial analysis with speech-based emotion detection. (International Journal of Latest Technology in Engineering Management & Applied Science)

  • Explainable Network Improves Multimodal Emotion Recognition — A proposed hybrid CNN-Transformer network integrates visual, audio, and textual information to enhance the transparency and accuracy of emotion recognition. (Asian Journal of Research in Computer Science)

  • Web Application Automates Dysarthria Speech Assessment — A new web-based tool integrates automatic speech recognition and acoustic analysis to assist speech-language pathologists in evaluating motor speech disorders. (JMIR Formative Research)

  • Voice-Controlled Assistance Systems Support Emergency Surgeons — Smart voice-controlled hands-free systems enable surgical teams to interact with medical devices and retrieve patient data without physical contact. (Journal of Nursing Future Care AI and Innovation)

  • Voice Agent Tested in Emergency Department — A case study evaluated ED GOAL-AI, a voice-based conversational agent designed to conduct structured values discussions with older adults in emergency departments. (doi.org)

  • Voice-Enabled Healthcare Platform in India — Swasth AI allows users to describe symptoms via voice in seven Indian languages to receive preliminary health assessments and book doctor appointments. (International Journal of Innovative Research in Computer and Communication Engineering)

  • AI Voice Cloning in Medicine — A randomized controlled trial evaluated the pedagogical value and student perception of AI-based voice cloning compared to traditional audio recording for prerecorded medical courses. (JMIR Medical Education)

  • Generative AI Framework Enhances Real-Time Translation — The proposed GAITA framework integrates speech-to-text, translation, and text-to-speech technologies to address accessibility and language barriers in India. (International Scientific Journal of Engineering and Management)

  • Nigerian Broadcasters Evaluate Text-To-Speech Technology — A study of Edo Broadcasting Service and Independent Television assessed how broadcasters perceive and use artificial intelligence for text-to-speech and language translation. (Zenodo (CERN European Organization for Nuclear Research))

  • Dari Speech Recognition Systems Evaluated — Researchers developed an isolated Dari words corpus to train and evaluate deep neural networks for low-resource automatic speech recognition. (Zenodo (CERN European Organization for Nuclear Research))

  • Real-Time Sign Language to Speech — Researchers built SignBridge AI, a webcam-based system that translates sign language gestures into spoken words with custom voice profiles. (Zenodo (CERN European Organization for Nuclear Research))

  • Voice-First Assistant for Indian MSMEs — Vyapar AI is an offline-capable mobile application that uses Hinglish voice commands and on-device speech recognition to help small businesses manage operations. (Zenodo (CERN European Organization for Nuclear Research))

  • Bilingual English-Marathi Speech Synthesis System — Researchers developed an end-to-end neural text-to-speech system based on Tacotron 2 that utilizes guided attention and a speaker encoder for voice cloning. (International Journal of Advanced Research in Science Communication and Technology)

  • Bilingual Silent Speech Command Dataset — The new TESSCCo dataset contains electroencephalography signals recorded during overt and covert speech commands in English and Spanish. (Scientific Data)

  • AI App Teaches Vocabulary To Multilingual Children — An AI-enabled web application utilizes automatic speech recognition and text-to-speech to deliver conversational vocabulary instruction to elementary students. (doi.org)

  • WhiteTesseract Integrates Conversational AI With XR — The new system enables in-situ cultural heritage interpretation by combining spatial intelligence with conversational AI to adapt to visitor backgrounds. (Journal on Computing and Cultural Heritage)

  • Cochlear Implant Users Evaluate Emotional Prosody — Researchers conducted cross-sectional studies to determine the capacity of cochlear implant users to perceive and categorize emotional prosody in speech. (German Medical Science (German Research Foundation))

  • Noise Statistics Alter Midbrain Speech Encoding — A study using unanesthetized rabbits reveals that background noise statistics alter the neural representation of speech in a frequency- and modulation-specific manner. (bioRxiv (Cold Spring Harbor Laboratory))

  • Analyzing Speech-Based Human-AI Interaction — A new study proposes a conversation analysis framework to examine communication, cognition, and collaboration in speech-capable AI design systems. (Proceedings of the Design Society)

  • Designing Voice Agents For Attentional Vulnerabilities — A new study investigates how AI proactivity and interaction structure influence the user experience of voice-based conversational interactions. (Journal of the Association for Information Systems)

  • BPMN4ALL Enables Process Modeling Through Speech — A new software prototype allows domain experts to automatically generate and validate business process models using speech or text input. (Journal of the Association for Information Systems)

  • Hybrid AI Framework Analyzes Live Stream Commerce — Researchers developed an analytics framework that uses speech-to-text transcripts and multimodal data to identify dynamic participant roles in live streams. (Journal of the Association for Information Systems)

  • AI Voice Replication Transforms Music Industry — The rapid integration of generative AI to replicate human voices and automate music production raises critical questions regarding authorship and intellectual property rights. (International Journal of Leading Research Publication)

  • Maudsley Hosts UK And Ireland Speech Workshop — Over 150 researchers and industry professionals gathered at King's College London to discuss speech technologies. (King's College London)

  • Voice Technology Reshapes Modern Customer Service — Customer support centers are increasingly adopting voice-powered conversational systems to interact with callers before they reach human representatives. (intlbm)

  • Smartphone Adoption Drives Voice Recognition Growth — Rising global demand and smartphone adoption are fueling significant opportunities in the voice and speech recognition software market. (GlobeNewswire)

  • Emotion Becomes Trust Signal In Voice — Voice AI is increasingly utilizing emotion as a trust signal, prompting new discussions around privacy, bias, and identity verification. (Biometric Update)

  • Odabashian Outlines Voice AI Role In Healthcare — Roupen Odabashian argues that voice AI's primary value in healthcare lies in managing administrative tasks outside of direct clinical interactions. (Oncodaily)

  • Taylor Swift Files Trademark For Her Voice — The singer filed to trademark her voice and likeness to protect herself against AI-generated deepfakes. (AOL.com)

  • Harvey Keitel Warns Against AI Voice Licensing — Actor Harvey Keitel issued a warning regarding the licensing of actors' voices for artificial intelligence. (Let's Data Science)

  • Celebrity Voice Cloning Sparks Controversy In China — Cheap AI services that clone celebrity voices, expressions, and lip movements for fake ads have prompted lawsuits from top stars. (아시아경제)

  • AI Voice Used In School Bomb Threat — Law enforcement is investigating a school bomb threat in Nebraska that was issued using an AI-generated voice. (KSNB)

  • AI Voice Cloning Drives Massive Losses — A Gallup survey of over 5,000 adults links AI voice cloning and deepfakes to $68 billion in U.S. fraud losses in 2025. (Gadget Review)

  • AI Voice Cloning Scams Target Canadians — Scammers are increasingly cloning the voices of family members to deceive and steal money from victims in Canada. (Money.ca)

  • AI Voice Cloning Fuels British Scams — New research shows one in three Brits feel unsafe as AI voice cloning drives a rise in fraudulent phone calls. (Verge Magazine)

  • Poway Man Loses Savings To Voice Scam — A Poway resident lost his family's life savings after scammers used AI to mimic his son's voice in a phone scam. (FOX 5 San Diego & KUSI News)

  • AI Voice Scams Cost Spanish Victims — Fraudulent phone calls leveraging AI voice cloning have cost Spanish victims an average of 530 euros. (APD Noticies)

  • Cyber Criminals Adopt Face And Voice Cloning — Cyber criminals in Hyderabad are increasingly using voice imitation and face-voice cloning tactics to steal identities and deceive victims. (Bhaskar English)

  • AI Voice Cloning Lowers Fraud Barriers — The rise of AI-generated voices has dramatically changed the economics of fraud by eliminating the need for skilled human impersonators. (Unite.AI)