Hugging Face Daily PapersChengqian Ma, Wenhao Feng, Weixuan Jin1 min readpaperintermediate
Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue
Summary
This paper introduces Duplex-MPE, a new benchmark for evaluating full-duplex AI assistants in multi-party conversations, focusing on when to speak, remain silent, or stop. It found that MiniCPM-o 4.5 performed best among tested open-weight systems, while others struggled with appropriate silence and accuracy.
- Duplex-MPE evaluates full-duplex AI assistants in 3-4 human + 1 assistant scenarios.
- It measures fresh response initiation, answer accuracy, silence preservation, and stopping when a human resolves a request.
- Models receive continuous audio without transcripts or turn boundaries, simulating real-time interaction.
- MiniCPM-o 4.5 leads on three scored capabilities among the five open-weight systems tested.
Developers of full-duplex AI assistants can use this benchmark to improve models for more natural and context-aware interaction in complex multi-party dialogues.
7/10

