• Yelp Introduces New AI Phone Assistant — Yelp has expanded its platform capabilities by adding an AI-powered phone assistant to help users. (forbes.com)

  • Voice AI Platform Simulates Realistic Interviews — Researchers proposed an AI-driven real-time voice-based interview platform that uses speech-to-text and text-to-speech technologies to enable natural conversational practice. (International Journal of Innovative Science and Research Technology (IJISRT))

  • Telugu Voice Assistant Aids Rural Farmers — A newly proposed intelligent farming assistant utilizes speech recognition to allow rural farmers to communicate naturally and receive crop recommendations in Telugu. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research))

  • Deepfake Audio Detection Framework Achieves High Accuracy — A new study leverages pretrained YAMNet embeddings and deep neural networks to classify deepfake audio attacks with up to 99.24% accuracy. (International Journal of Advances in Data and Information Systems)

  • Generative Framework Reconstructs Folk Music Vocals — Researchers proposed the Authenticity-Constrained Generative Framework to balance historical accuracy and creative flexibility in AI-generated vocal reconstructions of cultural heritage. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research))

  • Indonesian Copyright Law Reviews Voice Cloning — A legal study analyzed the criminal implications of unauthorized commercial voice cloning under Article 113 of Law Number 28 of 2014 concerning Copyright. (Jurnal Ragam Pengabdian)

  • Sign Language Translation App Integrates Speech — A new real-time translation application uses MediaPipe hand geometry and text-to-speech technology to convert Indonesian sign language into stable audio output. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research))

  • Speech Recognition Diagnoses Language Learner Intelligibility — A study of Turkish EFL learners utilized automatic speech recognition as a diagnostic tool to analyze word-level pronunciation errors and intelligibility rates. (Ahi Evran Üniversitesi Sosyal Bilimler Enstitüsü Dergisi)

  • Google Speech Commands Dataset Receives Extension — Researchers reconstructed and validated state-of-the-art keyword-spotting models using the original Google Speech Commands dataset alongside a newly developed extension from Wrocław University of Science and Technology. (International Journal of Electronics and Telecommunications)

  • Microwave Photonic Radar Enables Speech Recognition — A newly proposed broadband tunable microwave photonic radar system detects subtle skin displacements to simultaneously monitor vital signs and perform high-precision speech recognition. (Advanced Photonics Nexus)

  • Cross-Modal Fusion Enables Real-Time Emotion Detection — A reliability-aware gated cross-modal fusion framework was developed to achieve real-time audio-visual emotion recognition using facial expressions and speech cues. (Journal of Intelligent Decision Making and Information Science)

  • Auditory Rehabilitation Framework Integrates Whisper ASR — The proposed Beyond Hearing framework combines OpenAI's Whisper speech recognition with visual tracking to support self-directed auditory rehabilitation. (Audiology and Speech Research)

  • Voice Commands Drive Real-Time VR Creation — The "One Word, One World" system integrates speech recognition with generative AI to allow users to initiate and manipulate 3D assets in collaborative virtual environments. (Virtual Reality)

  • Voice Recognition Powers Smart Home Automation — A newly designed home automation system uses an Android application to convert voice commands into text, enabling elderly and disabled individuals to control household appliances via Bluetooth. (Uluslararası mühendislik araştırma ve geliştirme dergisi)

  • Neural Speech Synthesis Enhances Student Audiobooks — A systematic review explored the adoption of AI-enhanced audiobook systems featuring neural text-to-speech synthesis to support inclusive learning for differently-abled university students in Sri Lanka. (Journal of Research in Music)

  • Unified Framework Improves Multi-Speaker Conversation Analysis — A new research project aims to develop a robust system that combines speech recognition, speaker diarization, and topic segmentation to better understand complex, overlapping conversations. (California Digital Library)

  • Semantic Communication Optimizes Real-Time VoIP Systems — Researchers developed a dynamic semantic prioritization and real-time voice reconstruction framework to prevent perceptual degradation in latency-sensitive 6G VoIP services. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research))

  • Closed-Loop Controller Optimizes AI Voice Calling — A new study proposes a receding-horizon integer-programming controller to dynamically allocate outbound calls and reduce customer hold times in conversational AI voice-calling systems. (International Journal of Advanced Artificial Intelligence Research)

  • Tether Evo Decodes Brain Activity Into Speech — Researchers demonstrated a single model trained with an alignment technique that can generalize speech decoding across different people for brain-computer interfaces. (techcrunch.com)

  • Smallest.ai Raises Thirteen Million For Voice AI — The startup secured funding to develop a small, specialized voice model designed to mimic human conversation by listening, thinking, and speaking simultaneously. (techcrunch.com)

  • PolyAI Releases Audio Native Dialog Model — The company launched Dialog-RSN-1, a model that processes raw audio without automatic speech recognition to achieve sub-300 ms response latency. (scouts.yutori.com)

  • Telugu Speech Recognition Models Evaluated in Study — Researchers investigated the effectiveness of pre-trained Wav2Vec XLSR-53 and Whisper-Small models for automatic speech recognition in the Telugu language. (Frontiers in Artificial Intelligence)

  • Acoustic Similarities Impact Speech Emotion Recognition — A new study explores how cross-linguistic acoustic-phonetic similarities affect multilingual speech emotion recognition across languages like Bangla, Hindi, Odia, English, and German. (AIUB Journal of Science and Engineering (AJSE))

  • Bilingual Clinical Language Processing Framework Developed — Researchers created a patient-focused healthcare system combining speech-to-text and text-to-speech modules for English-Yoruba translation. (Nature Journal of Emerging Sciences Technologies and Innovations)

  • Real Time Speech Translation System Built — A new system utilizes Whisper-Speech-to-Text and MMS-Text-to-Speech APIs alongside voice activity detection to translate spoken language into English speech. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research))

  • Speech To Text Conversion Techniques Reviewed — A review paper evaluates traditional and deep learning approaches used to convert spoken language into written text. (African Journal Of Applied Research)

  • Paralinguistic Information Flow in Adapters Diagnosed — Researchers tested whether frozen speech-to-LLM adapters preserve sentence stress beyond transcripts and how the text LLM utilizes it. (Big Data and Cognitive Computing)

  • OpenAI Launches Presence For Voice Agents — The new offering supports real-time experiences across voice and chat agents to handle workflows like customer support and outbound sales. (openai.com)

  • New Framework Detects Adversarial Audio Attacks — Researchers proposed CAFAD, a plug-and-play detection framework designed to protect automatic speech recognition systems from adversarial examples. (Cybersecurity)

  • Apple May Charge For Siri AI Features — CEO Tim Cook hinted that Apple is developing a plan to address high computing costs by potentially charging heavy users of its revamped voice assistant. (axios.com, techcrunch.com)

  • Friend Wearable Upgraded With Unique Voice — The AI-powered companion device has been overhauled with a built-in speaker to project a consistent personality and combat loneliness. (techcrunch.com)

  • ElevenLabs Launches AI Powered Speech To Text — The company introduced software designed to transcribe audio and video content across more than 90 languages with human-like precision. (elevenlabs.io)

  • Voice AI Market Projected For Massive Growth — The global Voice AI market is expected to reach $98.4 billion by 2034 as enterprises transition toward hyper-autonomous, full-duplex systems. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research))

  • Brain Computer Interfaces Restore Lost Speech — A comprehensive review highlights advances in high-speed neural decoding systems capable of translating thoughts into words at 62 words per minute. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research))

  • Open Source TTS Evaluation Framework Introduced — VoxMetrix has been launched as the first open-source web framework implementing both MOS and EyeTrackingMOS metrics to compare synthetic audio naturalness. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research))

  • OpenAI Releases Two New Transcription Models — The company introduced GPT-Live-Transcribe and GPT-Transcribe in its API, showing significant semantic accuracy improvements on its context-aware benchmark. (community.openai.com)

  • Cartesia Optimizes Sonic Model For Real Time — The company is optimizing its Sonic-3.5 text-to-speech model to deliver fast, clear, and natural real-time conversations. (cartesia.ai)

  • Comparative Machine Learning Framework for Speaker Identification — Researchers evaluated machine learning models using Mel Spectrogram features from the VoxCeleb dataset, finding that SVM and CNN classifiers deliver competitive speaker identification performance. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research))

  • Density Ratio Approach Integrates Multiple Japanese ASR Models — A study demonstrates robust automatic speech recognition in unknown target domains by integrating multiple end-to-end models using a density ratio approach. (APSIPA Transactions on Signal and Information Processing)

  • Noise-Robust Speech System Built for Construction Environments — Researchers proposed a noise-invariant human-robot interaction system combining a domain-specific speech recognition agent and a vision-language model for robot control. (Journal of Computing in Civil Engineering)

  • Memristive Nanowire Networks Enable Neuromorphic Audio Feature Extraction — A simulated study demonstrates that memristive nanowire networks can extract compact features directly from raw audio for low-latency, low-power spoken-digit classification. (APL Machine Learning)

  • Encore AI Raises Thirty Million Dollars for Conversational Agents — Encore AI secured $30 million in Series A funding to analyze customer interactions and train autonomous voice agents for support and sales teams. (techcrunch.com)

  • Fish Audio Launches S2.1 Pro Alongside Seed Funding — Fish Audio raised a $52 million seed round and launched its S2.1 Pro voice model, offering an expressive, low-latency, multilingual stack with open-weight availability. (scouts.yutori.com, techcrunch.com)

  • Acoustic-Based Hate Detection System Developed for Arabic Dialects — A study proposed a lightweight hybrid GRU-LSTM network called STRUSS to detect hate speech in spoken dialectal Arabic using acoustic representations. (International journal of intelligent engineering and systems)

  • OpenAI Integrates GPT-Live Duplex Voice Control Into Developer Workflows — OpenAI has brought its full-duplex GPT-Live audio model directly into developer workflows on Codex and ChatGPT for desktop. (venturebeat.com)

  • Apple Plans Siri-Centered Smart Home Push with New Hub — Apple is preparing to launch a smart home hub device built around its new Siri AI assistant, alongside a refreshed HomePod mini and TV set-top box. (bloomberg.com)

  • SpaceXAI Teases Voice Uploads in Imagine Omni Video Generator — SpaceXAI previewed Imagine Omni, a synthetic video tool featuring character locking, multi-reference consistency, and built-in voice upload and recording. (scouts.yutori.com)

  • Multi-User Transformer Decoder Accelerates Speech Neuroprosthesis Calibration — A joint-user transformer model trained across six participants decoded cortical activity into text with high accuracy, reducing the training data needed for speech brain-computer interfaces. (bioRxiv (Cold Spring Harbor Laboratory))

  • VisionAssist App Uses Real-Time Text-to-Speech for Blind Users — A new smartphone-based application recognizes surrounding objects and converts text to speech using Microsoft's Edge TTS engine to assist visually impaired individuals. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research))

  • Text-to-Speech Technology Boosts L2 English Speaking Performance — A mixed-methods study revealed that text-to-speech technology significantly improves the English-speaking skills of secondary school students, with the most pronounced gains among low-achieving learners. (The JALT CALL Journal)

  • Large Language Models Trace Roots Back to ASR Error Correction — A new paper traces the genealogy of LLMs, arguing that their foundational architectures evolved from auxiliary tools originally designed to correct transcription errors in speech recognition systems. (WSEAS Transactions on Computers archive)

  • Machine Learning Framework Enhances Speech Emotion Recognition Accuracy — Researchers developed an automated speech emotion recognition framework that combines Convolutional Neural Networks with advanced feature selection techniques to capture complex vocal cues. (Zenodo (CERN European Organization for Nuclear Research), Zenodo (CERN European Organization for Nuclear Research), PeerJ Computer Science)

  • Continuous Wavelet Transform and ResNet Classify Imagined Speech — A new brain-computer interface method processes EEG signals using Continuous Wavelet Transform and a ResNet-50-D architecture to decode imagined speech without physical production. (IRO Journal on Sustainable Wireless Systems)

  • IoT System Uses Edge Processing for Efficient Speech Recognition — An energy-efficient IoT system built on the ESP32-CAM achieves 98.2% speech recognition accuracy while reducing transmitted data traffic by 92%. (Электросвязь)

  • Adaptive Speaking Task Implements LLM-Driven Spoken Dialogue Systems — Researchers proposed an adaptive, computer-delivered speaking assessment that uses large language models to evaluate responses and select follow-up questions in real time. (Language Testing)

  • Speech Disorder Detection Framework Combines Deep Learning and Optimization — A new hybrid model integrates attention-guided deep learning and metaheuristic optimization to improve local and global acoustic pattern representation in speech emotion recognition. (DÜMF Mühendislik Dergisi) Author Voice Cloning Sparks Publishing Industry Concerns — Synthetic voice tools are now capable of copying the voices of established authors to produce instant book narrations for social media. (ft.com)

  • AssemblyAI Releases Universal Three Point Five Pro — The new flagship speech-to-text model is designed to handle real-world audio with improvements in accuracy, latency, and language switching. (assemblyai.com)

  • AI Explored to Address Speech Therapy Shortage — Researchers are investigating whether artificial intelligence can expand access to speech therapy in Bangladesh while keeping trained clinicians central to care. (The Daily Star)

  • Vocalbeats AI Partners With AGI Playground Singapore — The Singapore-based AI audio company has signed on as a strategic partner for the upcoming event. (GlobeNewswire)

  • MaxMine Launches Voice Enabled Mining Fleet Assistant — The new MAXI assistant allows mining supervisors to query fleet data, generate shift notes, and receive alerts using voice. (IT Brief Australia)

  • Friend Relaunches Talking AI Pendant Wearable — The hardware startup has introduced a $249 screenless pendant equipped with a speaker and a monthly subscription fee for memory. (Startup Fortune, PYMNTS.com)

  • Google Redesigns Android Voice Search Interface — The homescreen voice search button is receiving a major update that merges AI Mode, Search Live, and song lookup. (9to5Google)

  • Workers Increasingly Talk to Computers via Voice — The adoption of AI voice tools like Gemini Spark is rising in offices, though the shift faces security warnings and European Union delays. (AD HOC NEWS)

  • GPT Live Voice Integrates SynthID Watermarks — OpenAI has embedded Google's audio watermarking technology into all ChatGPT Voice and API outputs ahead of EU AI Act enforcement. (Tech Times)

  • Alibaba Releases Qwen Audio Speech Recognition Model — The new Qwen-Audio-3.0-ASR-Flash model achieves over 95 percent accuracy on medical vocabulary and holds the lowest error rate globally. (AIBase)

  • Building Voice-Controlled AI Agents Simplified — A new guide breaks down the technical pipeline of voice-controlled AI agents into streaming speech recognition and turn detection. (KDnuggets)

  • Parlance Enhances Healthcare Contact Centers — Parlance has updated its healthcare voice AI platform to offer self-service configuration and AI-briefed agent transfers. (PR Newswire)

  • ValidSoft Unveils AI Trust Intelligence Stack — ValidSoft has launched a comprehensive voice identity framework to secure humans and AI agents through voice biometrics and synthetic audio detection. (AiThority)

  • SalesCloser Secures Major Social Media Client — SalesCloser has deployed its autonomous AI sales agents to handle high-volume applicant qualification for a top-five global social media platform. (GlobeNewswire)

  • Voice AI Solves Indian Hospital Phone Issues — Mid-sized hospitals in India are increasingly adopting voice AI to address severe communication and phone infrastructure gaps. (Nasscom)

  • Palabra.ai Leads Text-To-Speech Speed Benchmark — Palabra.ai achieved the top spot for latency on Coval's open-source text-to-speech benchmark with a record speed of 104 milliseconds. (StreetInsider)

  • Smallest.ai Raises Twenty-One Million Dollars — Smallest.ai secured a $13 million Series A funding round led by Seligman Ventures to develop its asynchronous Voice 4.0 and Hydra platforms. (The Tech Buzz, citybiz, Межа. Новини України., Porterville Recorder, PR Newswire, SiliconANGLE)

  • Journalist Tests Voice Cloning on Family — A Cybernews writer successfully cloned their own voice using AI to test if their mother could detect the scam. (Cybernews)

  • Tysa Celebrates One Year of Conversations — Conectys' multilingual AI voice agent, Tysa, completed its first year of enabling scalable, real-time customer experience interactions. (EIN News)

  • Senate Warns of Rising AI Scams — A bipartisan U.S. Senate panel warned that artificial intelligence and voice cloning are making financial scams targeting seniors significantly more convincing. (INDIA New England News, The Economic Times)

  • OpenAI GPT Transcribe Lowers Audio Costs — OpenAI's GPT Transcribe reduces audio transcription costs for developers by offering live speech tools, context, and keyword features. (Memeburn)

  • Rime Raises Twenty-Four Million Dollars — San Francisco-based speech AI developer Rime raised a $24 million Series A round to expand its text-to-speech models. (Slator)

  • xAI Launches Grok Voice Model Upgrade — xAI released Grok Voice Think Fast 2.0, a next-generation speech-to-speech model featuring improved transcription accuracy and faster inference speeds. (BigGo Finance, The American Bazaar, X.ai, GIGAZINE)

  • Hoocs.ai Launches High-Speed Audio Transcription — Hoocs.ai introduced a fast, cost-efficient AI audio-to-text converter designed to transform professional workflows. (EIN Presswire)

  • Heytruffle Secures Funding for AI Concierge — The company formerly known as RestoHost secured funding from Preface Ventures to expand its managed AI phone concierge services. (Trend Hunter)

  • Avatarin Builds Retail Agent with OpenAI — Retailer Yamada Denki deployed a 24/7 multilingual support agent built on OpenAI's GPT-Realtime to assist shoppers. (OpenAI)

  • PolyAI Releases Real-Time Voice Model — PolyAI launched Dialog-RSN-1, a voice dialog model that reasons directly over raw call audio to reduce latency and improve emotional awareness. (Tech Times, 디지털투데이, SiliconANGLE, CMSWire)

  • Soracom Launches Air RTC Gateway Service — Soracom introduced a cloud service that routes cellular voice from IoT devices directly to contact centers and AI agents based on SIM identity. (Cyprus Shipping News, The Fast Mode)

  • Swivl Releases Self-Storage Voice Agent Data — Swivl published data from the first quarter of 2026 detailing the performance and usage of its AI voice agents in self-storage operations. (Inside Self-Storage)

  • WellSpan Health Partners with Hippocratic AI — WellSpan Health expanded its partnership with Hippocratic AI to deploy generative voice AI agents across inpatient and ambulatory workflows. (HIT Consultant)

  • Nabla Launches Medical Dictation for Apple — Nabla introduced a medical-grade dictation product built specifically for Apple devices to advance its voice-first healthcare vision. (Morningstar)

  • Taco Bell Expands Voice AI Ordering — Taco Bell expanded its voice AI technology to nearly 900 U.S. restaurants in partnership with its voice AI provider. (Food On Demand)

  • Brothers Launch Savi Scam Protection App — Following a voice-cloning scam attempt on their mother, two brothers developed and launched Savi, an AI-powered scam protection application. (cbs8.com)

  • Advocates Push FTC to Probe Voice Cloning — Consumer advocates are urging the Federal Trade Commission to investigate an unnamed AI company's voice cloning product. (Consumer Federation of America)

  • MUSC Health Expands Emily Voice Agent — MUSC Health integrated SoundHound AI's voice platform "Emily" into retail and specialty pharmacy workflows with native Epic EHR support. (HIT Consultant)

  • Windows Eleven Update Upgrades Voice Typing — Microsoft released update KB5101681 for Windows 11, introducing upgrades to its built-in voice typing feature. (Ubergizmo)

  • Encore AI Raises Thirty Million Dollars — Conversational AI startup Encore AI secured $30 million to train client-specific voice agents and deepen CRM integrations for financial firms. (Межа. Новини України.)

  • Tata Communications Launches SMB Voice AI — Tata Communications and TTBS introduced a new voice AI platform built on Commotion's capabilities to support Indian small and medium businesses. (Inc42)

  • Vellum Adds Voice Mode to Assistant — Vellum introduced a new Voice Mode that processes spoken conversations through the same agent loop powering its text assistant. (TestingCatalog AI News)

  • Fish Audio Raises Fifty-Two Million Dollars — Fish Audio secured $52 million in seed funding and launched its S2.1 Pro voice model to compete in the expressive voice AI market. (Pulse 2.0, Tech Times, Startup Fortune)

  • Cencori and Spitch Partner in Africa — Cencori and Spitch partnered to deliver localized African voice AI models, including Yoruba, Hausa, Igbo, and Amharic, through a single API. (iAfrica.com)

  • New Orleans Tests AI for Nine-One-One — The city of New Orleans is testing artificial intelligence voice agents to answer non-emergency and emergency 911 calls. (New Orleans CityBusiness)

  • LG Gram Laptops Integrate Offline Dictation — LG Electronics integrated ActionPower's on-device voice recognition technology into its Gram laptop line to enable offline dictation. (BigGo Finance)

  • Krafton Releases Open Source Speech Model — South Korean game developer Krafton released its new audio foundation model, A.X K2 Raon-Speech, as open source on Hugging Face. (매일경제, The Korea Herald, Chosunbiz)

  • Google Adds Voice Dictation to macOS — Google rolled out new voice features for the Gemini app on macOS, allowing users to transcribe, edit, and summarize spoken requests. (blog.google, AppleInsider, jetstream.blog)

  • OpenHome and ElevenLabs Partner in Japan — OpenHome and ElevenLabs Japan launched a developer program offering ElevenLabs credits for voice AI projects built on OpenHome hardware. (WBOC TV, Yahoo Finance, Business Wire, The Daily Tribune News, 01net)

  • Exaforce Launches Voice-Powered Security App — Exaforce introduced ExaGo, a hands-free, voice-powered mobile application designed for security operations center teams. (Business Wire)

  • CallRail Expands Voice Assist Availability — CallRail made its Voice Assist solution available as a standalone product to all businesses while adding contextual AI texting. (citybiz)

  • Parlance Releases Parlance Twelve Platform Upgrade — Conversational voice AI provider Parlance launched Parlance 12 as its platform nears a milestone of two billion handled calls. (PR Newswire)

  • United Telecoms Launches Voice Agent Platform — South African unified communications provider United Telecoms introduced a proprietary AI voice agent platform for local businesses. (Telecompaper)

  • Cloudonix Contributes Call Transfer to Dograh — Cloudonix enabled native call transfers between AI voice agents and business phone systems by contributing code to the open-source Dograh project. (AiThority, EIN News)

  • Researchers Establish Vocal Biomarker Standards — Experts have established standards for vocal biomarkers to help predict diseases through voice recordings and AI analysis. (Medical Xpress)

  • La Mesa Police Adopt Voice Assistant — The La Mesa Police Department implemented an AI-powered voice assistant to handle non-emergency phone calls. (cbs8.com)

  • Yelp Host Integrates with OpenTable — Yelp Host expanded its voice AI capabilities by integrating with OpenTable to automatically manage restaurant reservations and takeout orders. (AOL.com, Yelp)

  • DXC Technology Partners with ElevenLabs — DXC Technology formed a strategic partnership with ElevenLabs to scale enterprise AI and voice innovation across its business. (CRN, PR Newswire)

  • AI Voice Phishing Scams Target Elderly Citizens — Fraudsters in the United States are increasingly using voice cloning technology that requires as little as three seconds of audio to impersonate family members and target older adults. (디지털투데이, GIGAZINE)

  • New Orleans Implements AI Emergency Dispatch Agents — The city of New Orleans is deploying artificial intelligence agents to answer 911 calls instead of human dispatchers. (Shreveport Times)

  • Nabla Launches Medical Dictation For Apple Devices — Nabla has released a medical-grade dictation product built specifically for Apple devices to expand voice-first AI technology beyond patient encounters. (PR Newswire)

  • Krafton Releases New Foundation Voice AI Model — Krafton has launched its voice AI foundation model, "A.X K2 Raon-Speech," on Hugging Face, achieving top performance rankings in Korean and English. (Seoul Economic Daily, 아시아경제, 디지털투데이)

  • Fish Audio Launches Multilingual S2.1 Pro Model — Fish Audio has released its S2.1 Pro production voice model, which supports 83 languages and focuses on real-time conversational speech. (TestingCatalog AI News)

  • Fish Audio Secures Fifty-Two Million Dollar Seed Funding — Voice AI startup Fish Audio has raised $52 million in seed funding to expand its real-time text-to-speech, voice cloning, and voice agent platform. (PR Newswire, The American Bazaar, SiliconANGLE)

  • APAC Voice Agents Struggle With Regional Accents — While AI voice agents are expanding rapidly across the Asia-Pacific region, their ability to recognize diverse regional accents, dialects, and code-switching remains limited. (TNGlobal)

  • ElevenLabs Launches Grants Program For Voice Startups — ElevenLabs has introduced a grants program offering early-stage startups 12 months of free access to its full conversational AI, text-to-speech, and speech-to-text platform. (elevenlabs.io)

  • Deterministic AI Secures Voice Communication Channels — Implementing deterministic AI is becoming critical to securing the voice communications path as artificial intelligence technologies rapidly advance. (Telecom Reseller / Technology Reseller News)

  • Rime Raises Twenty-Four Million Dollar Series A — Voice AI startup Rime has secured $24 million in Series A funding to advance its speech-to-speech models. (AI Insider)

  • German Utility Pilot Proves Voice AI Success — A pilot project by rhenag demonstrates how voice AI agents can eliminate waiting times and provide reliable 24/7 customer service for energy providers. (Xpert.Digital - Konrad Wolfenstein)

  • AI Virtual Receptionists Transform Small Business Operations — Small businesses are increasingly adopting AI virtual receptionists to handle calls, integrate with CRMs, and provide bilingual support. (Technology Org)

  • SoundHound And Five9 Compete In Customer AI — Enterprises are rapidly adopting conversational AI and intelligent virtual agents, driving competition between customer engagement platforms like SoundHound and Five9. (TradingView)

  • SoundHound Stock Rises On LivePerson Partnership Expansion — SoundHound AI shares have traded up following bullish market sentiment surrounding its expanding AI voice partnerships, including its deal with LivePerson. (StocksToTrade)

  • CFA Urges Investigation Into Speechify Voice Cloning — The Consumer Federation of America has filed a complaint with the FTC and state attorneys general urging an investigation into Speechify for facilitating AI voice cloning impersonation scams. (Consumer Federation of America, Consumer Federation of America)

  • Anthropic Upgrades Claude AI Voice Mode Capabilities — Anthropic has released an upgraded voice mode for its Claude AI model family to make voice-based interactions more helpful for users. (AI Business)

  • Scammers Clone Police Voices In Pennsylvania Scam — Fraudsters in York County, Pennsylvania, are using AI voice cloning to impersonate local sheriff's deputies and demand immediate payment for fake warrants. (PennLive.com, GovTech, ABC27)

  • Actors Accuse Tech Firms Of Unauthorized Cloning — Multiple international actors have accused technology companies of cloning their likenesses and voices without consent to produce AI micro-dramas. (Source)

  • Japan Backs Civil Liability For Voice Cloning — A government panel in Japan has backed civil liability for the unauthorized AI use of public figures' voices under the right of publicity. (The Japan Times)

  • OpenAI Launches Real Time Voice For Enterprises — OpenAI has expanded its voice technology with the launch of GPT-Live for real-time task collaboration and Presence for deploying customer-facing voice agents. (Startup Fortune, OpenAI, SBS 뉴스)

  • Cybercrime Unit Retools After Voice Ransom Scam — A Missouri sheriff's office cybercrime task force is refocusing its strategy after investigating a chilling scam involving artificial intelligence voice cloning for ransom. (kobaran.com, KRCG)

  • Boson AI Introduces Higgs RealTime Speech Model — Boson AI founder Alex Smola is targeting the voice AI market with the Higgs RealTime speech-to-speech model, featuring emotion control and support for over 100 languages. (Crypto Briefing, Fortune, The Cryptonomist)

  • Video Transcriber AI Enhances Transcription Platform Capabilities — Video Transcriber AI has expanded its platform with new features including video-to-text transcription, YouTube transcript generation, and an audio-to-text converter. (FinancialContent)

  • LEO Technologies Launches Verus Voice AI Biometrics — LEO Technologies has introduced Verus Voice AI, a voice-biometric tool designed to help corrections agencies detect and stop personal identification number abuse in real time. (AI Magazine, Correctional News)

  • Deepgram Enhances Amazon SageMaker AI Security Integration — Deepgram has updated its support for Amazon SageMaker AI by integrating AWS IAM Temporary Delegation to secure speech-to-text workflows. (Amazon Web Services (AWS))

  • Indian Real Estate Adopts AI Voice Agents — Real estate companies in India are increasingly deploying AI voice agents to handle the high volume of phone calls required during the house-hunting process. (Tech in Asia)

  • AI Voice Agents Face Infrastructure Challenges In India — Although AI voice agents are becoming critical infrastructure in India, they frequently struggle with local linguistic and environmental complexities. (Hindustan Times)

  • AI Vocal Remover Tools Gain Creator Popularity — Content creators, musicians, and educators are increasingly utilizing AI vocal remover and voice isolator tools to clean audio clips and lessons. (Breaking AC News)

  • Queensland Startup Replicates Voices For Cancer Patients — Australian startup Laronix is using artificial intelligence to replicate and restore the natural voices of patients who have lost their ability to speak due to disease. (SmartCompany)

  • FSU Researchers Warn Of AI Voice Scams — Researchers at Florida State University warn that AI-powered voice cloning combined with traditional tactics is making financial fraud exceptionally difficult for older adults to detect. (KJCT)

  • Hybe Liquidates AI Voice Startup Supertone Operations — South Korean entertainment giant Hybe is liquidating its AI audio company Supertone just three years after acquiring it for approximately $33.4 million. (BigGo Finance)

  • Krisp Focuses On Securing Enterprise Voice AI — Krisp is addressing real-world deployment challenges by focusing on noise cancellation to secure and scale enterprise-grade voice AI. (CXOToday.com)

  • Technical Guide Explains Voice To Voice AI — A new technical guide details how modern voice-to-voice AI models manage speech encoding, audio tokens, streaming, turn-taking, and latency. (HackerNoon)

  • Taco Bell Expands Voice AI Drive Thrus — Taco Bell has expanded its voice-enabled AI drive-thru ordering system to more than 890 restaurants across 38 states in the U.S. (Trend Hunter)

  • ElevenLabs Introduces Prompt Based AI Voice Design — ElevenLabs has launched Voice Design, a new feature and API that allows users to generate unique, realistic, or character-based voices using text prompts. (elevenlabs.io)

  • Production Voice Stack Built For African Languages — A new production voice AI architecture combines fine-tuned speech recognition, translation, and streaming text-to-speech specifically for Igbo, Yoruba, and Hausa. (HackerNoon)

  • ESTsoft Leads Government AI Dubbing Project Again — ESTsoft will lead the South Korean government's K-FAST project for the second consecutive year to accelerate the globalization of K-content using AI dubbing. (벤처스퀘어)

  • Instagram Adds Voice Tools To Reels Feature — Instagram has introduced new audio tools to Reels, including speech-to-text capabilities and voice effects. (Mashable)

  • AI Speech To Text Tools Accelerate Writing — New AI speech-to-text tools are helping users who prefer speaking over typing to quickly draft lengthy documents, emails, and notes. (Trend Hunter)

  • Meta Temporarily Adjusts Smart Glasses Audio Feature — Meta has walked back limits on its smart glasses' Conversation Focus feature, allowing users to temporarily continue utilizing the audio capability. (Engadget)

  • AI Notetakers Raise Privacy Concerns Among Professionals — While AI notetakers offer quick meeting summaries and action items, some professionals are questioning their use due to data privacy concerns. (Jacksonville Journal-Courier)