proomt

Search

Search posts, papers, and topics

dialogue

RSS
  1. 1

    StepAudio 3 Realtime Technical Report

    StepAudio 3 Realtime is an audio‑language foundation model that runs a continuous listen‑converse‑think‑act loop. It introduces Deep Perception for rich acoustic cue extraction, Seamless Duplex for handling pauses/back‑channels, and a Think‑While‑Speaking mechanism that lets the model reason in parallel with speech output. On benchmarks it scores 73.0 macro avg on StepAudioChat, 90.6 on MMSU, 98.…

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. 2

    Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX

    Cross‑corpus study of gaze behavior in two collaborative dialogue datasets (MapTask, MUNDEX) shows that task‑aligned references correlate with more task‑directed, less partner‑directed gaze, lower entropy and fewer transitions. Temporal gaze features (MapTask) and raw proportion features (MUNDX) modestly improve grounding prediction over baselines, but effects are small and diminish when aggregat…

    Hugging Face Daily Papersarxiv.org1 minpaper