proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersChengqian Ma, Wenhao Feng, Weixuan Jin1 min readpaperintermediate

Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue

Summary

This paper introduces Duplex-MPE, a new benchmark for evaluating full-duplex AI assistants in multi-party conversations, focusing on when to speak, remain silent, or stop. It found that MiniCPM-o 4.5 performed best among tested open-weight systems, while others struggled with appropriate silence and accuracy.

  • Duplex-MPE evaluates full-duplex AI assistants in 3-4 human + 1 assistant scenarios.
  • It measures fresh response initiation, answer accuracy, silence preservation, and stopping when a human resolves a request.
  • Models receive continuous audio without transcripts or turn boundaries, simulating real-time interaction.
  • MiniCPM-o 4.5 leads on three scored capabilities among the five open-weight systems tested.

Developers of full-duplex AI assistants can use this benchmark to improve models for more natural and context-aware interaction in complex multi-party dialogues.

7/10

Related reading

  1. SteerDuplex: Steerable Duplex Speech Dialogue Models

    The paper presents SteerDuplex, a full‑duplex speech dialogue model that can be steered along tone, persona, and speed via instruction following, and introduces the SteerBench benchmark to evaluate such steerability. Supervised training yields a 44.5 % pass‑rate lift, and reinforcement‑learning fine‑tuning improves interruption handling and reduces pause barge‑ins, though reward hacking remains a…

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. Realtime-Venus: A full-duplex interaction system with asynchronous delegation

    Realtime-Venus is a full‑duplex, multimodal dialogue system built from two separately trained 9B models (Omni for audio‑visual, Audio for spoken interaction). It uses a shared causal timeline and a dual‑loop runtime that lets foreground interaction continue while a background Harness executes delegated tasks asynchronously. The paper reports benchmark scores where Realtime‑Venus‑Omni leads on six…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. How OpenAI Built GPT-Live

    OpenAI’s GPT‑Live‑1 is a full‑duplex voice model that can listen and speak simultaneously by tokenizing audio (including silence) and emitting tokens on an ~80 ms clock. By keeping the model small and fast and delegating heavy reasoning to a separate LLM, OpenAI solves the turn‑detector problem of earlier cascaded and turn‑based systems while meeting sub‑100 ms latency requirements. The post walk…

    ByteByteGobytebytego.com14 min
  4. OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

    OmniVChat defines native audio‑visual dialogue where a model consumes raw audio and video streams and replies in text. The authors build OmniVChat‑Studio, a multi‑agent simulator that generates single‑ and multi‑turn audio‑visual conversations, and use it to create OmniVChat‑Bench, a benchmark covering five dialogue abilities. They also propose OmniVChat‑RL, a reinforcement‑learning reward that b…

    Hugging Face Daily Papersarxiv.org2 minpaper
  5. Verifiable Social Reasoning for LLM Assistants

    The paper introduces Fuse, a multi‑agent simulation that gives LLM assistants a verifiable ground‑truth task for social reasoning by hiding a target agent’s motive and letting a user‑mediated conversation infer it. Experiments on 12 LLMs show user mediation makes reasoning harder, models are biased by user framing, need more detail than humans, and longer chats don’t always help.

    Hugging Face Daily Papersarxiv.org1 minpaper