1
Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change
The paper introduces VoxParity, a benchmark that swaps audio cues while keeping transcripts fixed to evaluate whether voice agents act on what they hear. Across 183 scenarios, most current systems fail to adjust their actions, exposing a strong bias toward text‑only reasoning.
Hugging Face Daily Papersarxiv.org1 minpaper
