proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersHaibo Wang, Jiteng Mu, Jialu Li1 min readpaperadvanced

OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

Summary

OmniSeek is a framework that turns a large language model into a multi‑turn audio‑visual reasoning agent by letting it decide when to look or listen and which temporal windows to fetch as evidence. The authors train it on a synthetic 170K trajectory dataset, then refine with reinforcement learning and an Audio‑Visual Necessity loss, reporting consistent gains on several benchmarks.

  • The agent dynamically selects when to query audio or visual data and the length of the temporal segment, appending raw evidence back into its context for subsequent reasoning.
  • Training uses a synthetic corpus (OmniTraj‑170K) of multi‑hop chain‑of‑thought trajectories with interleaved audio‑visual evidence, followed by supervised fine‑tuning.
  • A two‑stage RL stage with verifiable rewards further optimizes the evidence‑seeking policy.
  • An Audio‑Visual Necessity objective penalizes single‑modality shortcuts, encouraging true multimodal reasoning.

Teams building multimodal AI agents will find the tool‑integration protocol and training recipe useful for enabling efficient, evidence‑driven reasoning across audio and video.

6/10

Related reading

  1. OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

    OmniVChat defines native audio‑visual dialogue where a model consumes raw audio and video streams and replies in text. The authors build OmniVChat‑Studio, a multi‑agent simulator that generates single‑ and multi‑turn audio‑visual conversations, and use it to create OmniVChat‑Bench, a benchmark covering five dialogue abilities. They also propose OmniVChat‑RL, a reinforcement‑learning reward that b…

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

    OmniHarness introduces a symbolic‑policy framework that extracts reusable procedural knowledge from multimodal LLM‑driven visual generation runs. By decoupling task logic from instance inputs, the system can instantiate, adapt, and compose policies for new visual tasks, using intermediate verification for on‑the‑fly refinement while keeping the underlying model frozen. Self‑directed practice task…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

    HarnessVLN introduces a zero‑shot, training‑free embodied navigation framework that wraps a multimodal LLM in an "Agent Harness" – a tool‑based protocol that validates planner actions against spatial evidence, tracks progress with hierarchical event memory, and maintains a persistent spatiotemporal graph for recovery. The system works for instruction‑following and object‑goal tasks, achieving 60.…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    The paper pinpoints two weaknesses in latent‑reasoning pipelines—suboptimal visual token encoders and deterministic RL sampling—and proposes Scaffolding Minds, which trains a dedicated scaffolding encoder and a stochastic RL sampler (learning both mean and variance). This yields up to +19 points improvement on spatial planning tasks and +5.6 points average gain on nine visual reasoning benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

    The paper presents AV‑GRPO, a diffusion‑based reinforcement‑learning framework that treats audio and video generation as separate but coordinated tasks by anchoring rollouts to each modality and freezing the opposite tower during optimization. Evaluated on the new 5DAV dataset and benchmarks, it achieves higher fidelity, better text‑modality alignment, and tighter audio‑video sync than the previo…

    Hugging Face Daily Papersarxiv.org1 minpaper