Voice AI This Week: Bigger Models, Real Money, and a Fight Over a Dead Man's Voice

TL;DR. The model race got louder: Google shipped Gemini 3.5 Transcribe across 85+ languages on August 26, IBM claimed on August 25 that Granite Speech 5.0 transcribes 3.5 hours of audio per second, and a run of new models landed on August 27 alone, Daily's open-weights PhoneLLM, Tavus Sparrow-2 for conversational understanding, and Cartesia Sonic-3.6 for text-to-speech.
The plumbing got cheaper and more programmable too: LiveKit moved to per-second billing, ElevenLabs shipped a CLI, Retell launched in-call CRM Workflows, and AssemblyAI put a 4B rewrite model on its gateway. Real money moved, with Spear Street's Instinct raising a $250M Series B that lifts its total funding to $350M at a $2.5B valuation, and India's Ringg raising $10M from Peak XV on August 25. The dark side kept pace: a phishing-as-a-service kit now runs voice AI agents to steal iPhone passcodes, and on August 25 Carl Sagan's estate sued Luma AI over an ad that allegedly cloned his voice without permission. Our take: the infrastructure layer is consolidating and repricing fast, and consent law is the story to watch.

New models and launches

Google launched Gemini 3.5 Transcribe with automatic filler-word removal across 85+ languages (August 26, 2026). Google's new speech-to-text model strips "ums" and "ahs" and supports over 85 languages. Google says it streams in real time through a Live API with integrations for LiveKit, Pipecat, and Agora. (Google blogThe VergeArs Technica)

IBM says Granite Speech 5.0 transcribes 3.5 hours of English speech in one second (August 25, 2026). IBM released two compact English speech-recognition models as part of its Granite 4.2 family, claiming they process 3.5 hours of English audio per second. (IBM ResearchUnite.AI)

Daily open-sourced PhoneLLM Alpha 1, a small open-weights model for voice agents (August 27, 2026). The Pipecat team released an open-weights LLM tuned for low-latency, multi-turn calling. Daily positions it at a fraction of the cost and latency of larger general-purpose models with comparable performance on voice-agent tasks, from inbound customer service to outbound calling, and ships it alongside a companion phone-agent benchmark. Our take: open weights plus an outbound-calling focus is exactly the direction we back. (Daily / Pipecat)

Tavus shipped Sparrow-2, a real-time conversational understanding model (August 27, 2026). Sparrow-2 jointly models turn-taking, interruptions, backchannels, and the surrounding acoustic scene as a single streaming system, built for noisy real-world settings like cafés, airports, and retail floors. (Tavus)

Cartesia released Sonic-3.6 (August 27, 2026). The updated text-to-speech model expands multilingual coverage and, per Cartesia's own blind head-to-head tests across fifteen locales, was preferred by listeners over Sonic-3.5 in up to 93% of comparisons. Preference figures are the company's own numbers. (Cartesia)

Sesame published TurnBench, an open benchmark for conversational turn-taking (week of August 24, 2026). TurnBench hand-annotates end-of-turn and interruption events in dual-channel human conversations and ranks models on recall, false-positive rate, and latency, with a public leaderboard, dataset, and paper. We could not pin an exact announcement date from a secondary source, so it is dated to the week it surfaced. (Sesame TurnBench)

Infrastructure and tooling

LiveKit moved agent, recording, WebRTC, and SIP usage to per-second billing (August 24, 2026). Charges now meter per second with a ten-second minimum increment instead of by the minute, across all plans including enterprise contracts, with no customer-facing code changes. LiveKit says it is already live and shows up on September bills for August usage. (LiveKit)

Retell AI launched Workflows to pull CRM and helpdesk context into calls and write outcomes back(August 28, 2026). Retell Workflows connects Salesforce, HubSpot, Zendesk, Cal.com, and Calendly directly, so an agent loads context before it speaks and writes the result the moment the call ends, replacing a Zapier or n8n layer teams maintained themselves. (Retell AI)

ElevenLabs shipped CLI v1, bringing its API and agents-as-code to the terminal (August 24, 2026). Every operation is a documented subcommand with structured JSON output and a --dry-run preview mode, and workspace agents can be pulled into local config files and edited like code. (ElevenLabs)

AssemblyAI added Qwen3.5 4B to its LLM Gateway for fast voice-rewrite tasks (week of August 24, 2026). The hosted model targets dictation cleanup and transcript rewrites at roughly 600ms in a latency-optimized 32k-context configuration, priced at $0.10 input and $0.50 output per million tokens. Exact announcement date not independently pinned, so it is dated to the week it surfaced. (AssemblyAI)

The money

Spear Street's Instinct raised a $250M Series B, lifting total funding to $350M at a $2.5B valuation(announced August 26, 2026). The AI assistant, reachable by text and phone call and founded by 23-year-old Noah Shinn, closed a $250M Series B co-led by Index Ventures and Benchmark, bringing total funding to $350M. (Precision note: the $350M widely quoted is total funding, not the size of this round.) (TechCrunchQuartz)

India's Ringg raised $10M in a Series A extension led by Peak XV (August 25, 2026). The Bengaluru enterprise voice AI startup will use the funding to push beyond phone calls. (TechCrunchEntrackr)

HeyBreez raised a $2.5M seed round led by Lunara Partners for MENA voice AI infrastructure (week of August 22, 2026). The multilingual voice AI platform closed an oversubscribed seed round as monthly call volume passed one million. (FWDStartSlator)

Law, safety, and consent

Carl Sagan's estate sued Luma AI over an ad that allegedly cloned his voice (filed August 25, 2026). Druyan-Sagan Associates filed suit in California, alleging the startup used a cloned version of the late astronomer's voice in an advertisement without permission. (ForbesLaw360)

Australian voices are being cloned without consent, prompting calls for a crackdown (reported August 25, 2026). Unauthorized clones of Australian politicians, celebrities, and citizens have sparked demands for regulation. (7NEWS)

A phishing-as-a-service kit named AnonyMousKIT uses voice AI agents to steal iPhone passcodes(reported August 25, 2026). Researchers documented the platform automating passcode theft with voice agents. (BleepingComputer)

Scammers are using fake Apple support calls to extract iPhone passcodes (reported August 25, 2026). Criminals combine AI and fake support calls to trick owners of stolen iPhones. (Cybernews)

83% of AI voice-scam victims in India reported financial losses (reported August 23, 2026). A majority of victims lost money as the voice-fraud threat grew, per Media India Group. (Media India Group)

In the wild: products, enterprise, and culture

Plaud unveiled the Plaud One earbuds with an eSIM-enabled case (August 27, 2026). The $249.99 Explorer Edition wearable transcribes conversations and connects to productivity tools, shown ahead of IFA in Berlin. (TechCrunchThe Verge)

Equal AI screens calls in eight Indian languages and made the Forbes Asia 100 to Watch (August 24, 2026). The Hyderabad startup's assistant screens incoming calls with live transcription. (Forbes)

Grab Singapore launched AI Call-A-Ride for voice-based ride booking (August 25, 2026). The ride-hailing company added voice-command booking for accessibility. (Marketech APAC)

Suki launched a standalone AI clinical dictation tool with Epic and MEDITECH integration (August 25, 2026). The AI scribe unbundled dictation so health systems are not forced into an enterprise-wide documentation suite. (Business WireMobiHealthNews)

GC MediAI reports 97.1% accuracy transcribing consultation-room conversations (August 26, 2026). The company's "Doctor's Companion AI" logged a 97.1% accuracy rate turning consultations into electronic charts, presented at a media day on August 26. (Asiae)

CARIVA is building a multilingual hospital phone agent with emergency escalation (August 28, 2026). The Thai startup's real-time voice agent handles hospital phone lines and escalates potential emergencies. (OpenAI)

Filmmaker S.S. Rajamouli is keeping human dubbing over AI for emotional nuance (reported August 23, 2026). The director is prioritizing human voices to preserve the emotional nuance of translation. (Deccan Chronicle)

A father configured ChatGPT's voice mode into a tutor that quizzes his kids (reported August 24, 2026). The parent set up ChatGPT voice to quiz his children on spelling and math. (Business Insider)

Papers this week

The arXiv preprints below carry August 2026 identifiers. Each links to its primary source.

  • DuplexGen proposes decoupling content, timing, and acoustics to improve turn-taking in full-duplex speech synthesis. (arXiv)

  • VoiceMem introduces a streaming dual-brain memory architecture for voice assistants, evaluated on ChatMem-Bench. (arXiv)

  • A study finds ASR errors amplify in multi-hop RAG, with accented speakers suffering higher word error rates that degrade voice-assistant answers. (arXiv)

  • TurboBias 2.0 presents a context-biasing method that adjusts decoding scores for streaming ASR without retraining. (arXiv)

  • Pixel-TTS renders text as images to improve robustness and cross-lingual adaptation in speech synthesis. (arXiv)

  • A paper quantifies the gap between high ASR benchmark scores and real-world utility. (arXiv)

  • query-based multimodal framework combines acoustic features and ASR transcripts to detect Alzheimer's from speech. (Frontiers in Aging Neuroscience)

  • A clinical study shows speaker-verification-inspired ECAPA-TDNN embeddings can detect obstructive sleep apnea from daytime speech. (Scientific Reports)

  • SpectralTrojan demonstrates a psychoacoustically constrained backdoor attack on keyword spotting and speaker-identification models. (Scientific Reports)

  • A paper questions whether voices are unique biometric imprints by probing automatic speaker-recognition systems. (arXiv)

  • Research on Tibetan cross-lingual voice cloning finds that disabling prompt latent prefill mitigates dialect leakage. (Preprints.org)

  • PersonaVoice adapts F5-TTS with length-adaptive flow matching to clone a voice from roughly one second of reference audio. (Zenodo)