Hugging Face Daily PapersBhavik Mangla1 min readpaperadvanced
Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change
Summary
The paper introduces VoxParity, a benchmark that swaps audio cues while keeping transcripts fixed to evaluate whether voice agents act on what they hear. Across 183 scenarios, most current systems fail to adjust their actions, exposing a strong bias toward text‑only reasoning.
- VoxParity benchmark presents 183 audio‑varying scenarios across 14 sectors, keeping the transcript constant to test voice agents' reliance on acoustic cues.
- Only 11 of 23 evaluated systems pass the audio‑aware test, indicating most agents ignore critical non‑verbal information.
- Systems over‑react to routine requests even when audio signals emergencies (41% vs 12% error rates), showing a transcript bias.
- Providing explicit voice descriptions or rule text recovers part of the performance gap, but emotion cues remain largely missed.
Anyone building voice assistants, emergency‑call automation, or multimodal LLMs should care because current models still miss critical acoustic signals.
7/10
