How Voice AI Benchmarks Work: WER, TTS Arenas, Turn-Taking (2026)
Master voice AI evaluation by analyzing Word Error Rate, competitive TTS arenas, and turn-taking metrics to build advanced conversational systems.
Summary
Components of a Voice AI Benchmark
Every benchmark consists of three main components, similar to a driving test:
The test set: The specific audio or conversations used to test the models.
The metric: The resulting numerical score.
The conditions: Factors such as whether the audio is clean or from a phone line, whether it is streaming, and other specific settings.
Altering any of these three components can change the model rankings.
Speech-to-Text Benchmarks
The standard metric for speech-to-text is the Word Error Rate (WER), which calculates the sum of swapped, dropped, and added words divided by the total words spoken. A lower score is better. Public leaderboards, such as Hugging Face's Open ASR Leaderboard, report WER and model speed. However, these benchmarks often use test sets that are cleaner than real-world phone calls, and they frequently clean up text before scoring, meaning formatting differences (like how numbers are written) may not be reflected.
Text-to-Speech Benchmarks
Because there is no single correct answer for text-to-speech, evaluation relies on human judgment:
Mean Opinion Score (MOS): The traditional method where listeners rate audio clips on a scale of 1 to 5, and the scores are averaged.
Arenas: A newer approach (used by platforms like Artificial Analysis and TTS Arena) where users listen to the same phrase spoken by two anonymous models, select their preference, and the votes are compiled into chess-style ratings.
A limitation of these benchmarks is that a short, high-quality clip does not indicate whether a model can correctly pronounce specific product names or how quickly it delivers the initial audio.
Turn-Taking and Full Agent Benchmarks
Turn-taking benchmarks evaluate whether a model can correctly identify when a user has finished speaking. Responding too quickly cuts the user off, while responding too slowly creates dead air. LiveKit offers an open benchmark for turn-taking using real conversations in 14 languages, measuring cut-offs against wait times.
Full agent benchmarks, such as Tau-Voice (developed by Sierra and Princeton), test whether an agent can complete entire customer service tasks. These tests use simulated callers with various accents, background noise, and compressed phone lines, evaluating performance on a pass/fail basis by checking database changes. Currently, voice agents pass fewer tasks than text agents on the same assignments.
Evaluating Benchmarks
Automated benchmark scores do not always align with human preference; a model can score highly on a reasoning benchmark but still lose a listener vote. Before trusting a benchmark, you should ask five questions:
Who ran the benchmark, and do they sell one of the tested models?
What kind of audio was used (clean recordings or realistic calls)?
Which metric was used, and is a lower or higher score better?
What were the testing conditions (streaming vs. non-streaming, typical vs. worst-case latency)?
Can the benchmark be rerun using open data and code?
To find benchmarks, resources like voicebenchmarks.com and the Audio Benchmark Index on GitHub https://kennethli319.github.io/audio-benchmark-index/ categorize hundreds of options. Ultimately, the most reliable benchmark is your own, conducted by testing models on a sample of your own real customer calls.
Related Content

How TTS Models Are Trained: Data, Cleaning, Cloning (2026)

How Text-to-Speech (TTS) Works, Step by Step (2026)

Voice AI Stack Explained: Models, Frameworks & Platforms (2026)

How Voice Agents Work: STT, LLM, TTS, WebRTC & Speech-to-Speech

I Built a Minimalist AI Note Device

BreezeTTS2 - 100% Local Real-Time Voice

Voice @ AI Engineer

100% Local AI Speech to Speech with RAG - Low Latency | Mistral 7B, Faster Whisper ++