proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersBhavik Mangla1 min readpaperadvanced

Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change

Summary

The paper introduces VoxParity, a benchmark that swaps audio cues while keeping transcripts fixed to evaluate whether voice agents act on what they hear. Across 183 scenarios, most current systems fail to adjust their actions, exposing a strong bias toward text‑only reasoning.

  • VoxParity benchmark presents 183 audio‑varying scenarios across 14 sectors, keeping the transcript constant to test voice agents' reliance on acoustic cues.
  • Only 11 of 23 evaluated systems pass the audio‑aware test, indicating most agents ignore critical non‑verbal information.
  • Systems over‑react to routine requests even when audio signals emergencies (41% vs 12% error rates), showing a transcript bias.
  • Providing explicit voice descriptions or rule text recovers part of the performance gap, but emotion cues remain largely missed.

Anyone building voice assistants, emergency‑call automation, or multimodal LLMs should care because current models still miss critical acoustic signals.

7/10

Related reading

  1. SteerDuplex: Steerable Duplex Speech Dialogue Models

    The paper presents SteerDuplex, a full‑duplex speech dialogue model that can be steered along tone, persona, and speed via instruction following, and introduces the SteerBench benchmark to evaluate such steerability. Supervised training yields a 44.5 % pass‑rate lift, and reinforcement‑learning fine‑tuning improves interruption handling and reduces pause barge‑ins, though reward hacking remains a…

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

    OmniSeek is a framework that turns a large language model into a multi‑turn audio‑visual reasoning agent by letting it decide when to look or listen and which temporal windows to fetch as evidence. The authors train it on a synthetic 170K trajectory dataset, then refine with reinforcement learning and an Audio‑Visual Necessity loss, reporting consistent gains on several benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

    OmniVChat defines native audio‑visual dialogue where a model consumes raw audio and video streams and replies in text. The authors build OmniVChat‑Studio, a multi‑agent simulator that generates single‑ and multi‑turn audio‑visual conversations, and use it to create OmniVChat‑Bench, a benchmark covering five dialogue abilities. They also propose OmniVChat‑RL, a reinforcement‑learning reward that b…

    Hugging Face Daily Papersarxiv.org2 minpaper
  4. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.

    Hugging Face Daily Papersarxiv.org1 minpaper