Hugging Face Daily PapersYanjie Zhang, Nanchen Hu, Yushi Sun1 min readpaperadvanced
Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
Summary
The authors present VGBench, a 1,018‑item diagnostic suite that tests whether audio LLMs correctly mute actions when the acoustic context shifts (e.g., speaker switches). Training with VoxGate dramatically improves gating, raising mute rates from <15% to >90% while preserving tool‑call accuracy.
- VGBench isolates action‑level addressedness across side‑talk, self‑talk, and speaker‑switch scenarios with controlled acoustic variables.
- Six raw Audio LLMs rarely mute under speaker‑switch conditions, achieving at most a 14% mute rate.
- Supervised fine‑tuning (VoxGate) raises mute rates to 91.3% on switched commands without harming tool‑call performance on wearer speech.
- Factorized controls reveal that source‑change alone drives muting, while far‑field rendering effects vary by model.
Voice‑assistant developers need reliable gating to avoid unintended actions; this benchmark and training recipe provide a concrete way to evaluate and improve that behavior.
7/10

