Hugging Face Daily PapersHaibo Wang, Jiteng Mu, Jialu Li1 min readpaperadvanced
OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
Summary
OmniSeek is a framework that turns a large language model into a multi‑turn audio‑visual reasoning agent by letting it decide when to look or listen and which temporal windows to fetch as evidence. The authors train it on a synthetic 170K trajectory dataset, then refine with reinforcement learning and an Audio‑Visual Necessity loss, reporting consistent gains on several benchmarks.
- The agent dynamically selects when to query audio or visual data and the length of the temporal segment, appending raw evidence back into its context for subsequent reasoning.
- Training uses a synthetic corpus (OmniTraj‑170K) of multi‑hop chain‑of‑thought trajectories with interleaved audio‑visual evidence, followed by supervised fine‑tuning.
- A two‑stage RL stage with verifiable rewards further optimizes the evidence‑seeking policy.
- An Audio‑Visual Necessity objective penalizes single‑modality shortcuts, encouraging true multimodal reasoning.
Teams building multimodal AI agents will find the tool‑integration protocol and training recipe useful for enabling efficient, evidence‑driven reasoning across audio and video.
6/10