By Bhavik Mangla
A customer agrees to a utility disconnection: "Yes, I understand. Go ahead and schedule the disconnection for Friday, then." Behind the voice, a medical monitor is beeping. Asked what they heard, all four leading systems in a new benchmark identified the monitor. Three scheduled the disconnection anyway, and none applied the hold that utility rules require. Of all 28 systems tested, 20 scheduled the disconnection and none acted correctly.
VoxParity, a benchmark released on arXiv on 30 September, has 183 such scenarios. It tests whether an agent changes its action when the sound of a call changes the right one.
What it found
Errors run toward the words. All 28 systems, from 11 vendors and three serving modes, carry out the routine request on protective calls more often than they over-react on clean calls: 41% against 12% pooled (descriptive). A pipeline that only reads the words does so on 58%.

The median system does no better than that pipeline. It scores 0.34 on calls with an audible cue, the pipeline's own score. Only 11 of the 23 systems that can also run on a transcript act on the audio beyond what the words explain, including one of seven production realtime agents. Five of the 11 vendors have no system that does. The largest gains come from Qwen3.8-Omni and gemini-3.7-flash.
Heard, then not acted on. In exploratory analyses, most of the leading systems' misses are on cues they heard.
At the frontier, the loss is in deciding, not hearing (exploratory). For the four leading systems, perfect hearing would add 0.04 credit; perfect deciding would add 0.28.
For reference, the highest-credit system (0.57) is within a few points of volunteers who played the calls as a game (0.61).
How the test works

Each scenario holds one transcript fixed while the audio changes, and with it the correct typed tool call, following the sector's rules and practice. The executed call is scored without an AI judge. A system passes the words-only null test when hearing the call moves its actions more than it moves the pipeline (Whisper-large-v3-turbo, then gpt-oss-120b), Holm-corrected.
Run it on your agent
The harness and scorer are open source. Connect an agent through one Python class or an HTTP endpoint and run it on a public split of 40 scenarios (81 clips). The other 143 are held out with published hashes and evaluated on request.
Links
Leaderboard: https://bhavik-mangla.github.io/voxparity-bench/
Hugging Face Space: https://huggingface.co/spaces/bhavikmangla/voxparity-leaderboard
Written by Bhavik Mangla
About the author
Bhavik Mangla is a Senior Software Engineer at Zenarate, where he builds Evolve, the company's enterprise voice AI agent platform. He created VoxParity, the open benchmark behind this article. Before Zenarate he was an AI and data engineer at BlackRock, working on the Aladdin investment platform, and he founded Zoobi, a social commerce platform for Indian artisans. He is also an organization admin at the open-source group AOSSIE and has mentored for Google Summer of Code.
Connect with Bhavik on LinkedIn



