1
OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
OmniSeek is a framework that turns a large language model into a multi‑turn audio‑visual reasoning agent by letting it decide when to look or listen and which temporal windows to fetch as evidence. The authors train it on a synthetic 170K trajectory dataset, then refine with reinforcement learning and an Audio‑Visual Necessity loss, reporting consistent gains on several benchmarks.
Hugging Face Daily Papersarxiv.org1 minpaper
