proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJuyi Lin, Zhiqiang Lao, Jiali Cui1 min readpaperadvanced

LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

Summary

LEAP is a framework for long audio-video question answering that addresses context limits by retrieving relevant evidence blocks. It uses a two-stage process of block-wise localization and a bounded answer pass, achieving significant performance improvements on AVQA benchmarks.

  • LEAP decouples evidence localization from reasoning for hour-scale audio-video inputs.
  • It divides recordings into fixed blocks, applying a lightweight pass to score and select short candidate windows.
  • Only the highest-ranked windows are re-encoded for the final answer pass, keeping context size independent of recording duration.
  • Localization can use pre-computed transcripts, while the final answer pass uses raw audio-visual streams for fine-grained evidence.

Engineers developing multimodal AI systems for long-form audio-video content will find this relevant for improving efficiency and accuracy in question answering by managing context effectively.

8/10

Related reading

  1. OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

    OmniSeek is a framework that turns a large language model into a multi‑turn audio‑visual reasoning agent by letting it decide when to look or listen and which temporal windows to fetch as evidence. The authors train it on a synthetic 170K trajectory dataset, then refine with reinforcement learning and an Audio‑Visual Necessity loss, reporting consistent gains on several benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

    The authors present VākQA, a 2,001‑question spoken factoid QA benchmark for Telugu with audio, transcriptions, and human‑verified answers, and they validate automatic evaluation methods against human ratings. Using this setup they show that translation loses cultural nuance, ASR errors alter meaning, and cascaded ASR‑MT errors degrade model performance.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

    AnswerMap is a novel, training-free, black-box method for generating faithful spatial interpretability maps for Vision-Language Models (VLMs) directly from their output posteriors. It queries the VLM with image bands and yes/no relevance questions, demonstrating higher faithfulness than attention maps and enabling new applications like object localization and hallucination detection.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

    TRACE is a condition-aware benchmark for streaming video understanding that makes evidence timing and trigger conditions explicit, and measures answer quality, timeliness, workload, and reliability. Experiments on 1,240 records show that identical QA accuracy can hide large differences in completion, false alarms, and processing cost.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Native Action-Prior Learning from Videos for World Action Models

    Native Action-Prior Learning (NAVA‑WAM) trains robot action policies directly from observation‑only videos by matching future video flow through a joint attention mechanism, then fine‑tunes with a small set of labeled demos. Experiments show it beats prior methods on both in‑distribution and out‑of‑distribution tasks and transfers to real robots with fewer action labels.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

    OmniVChat defines native audio‑visual dialogue where a model consumes raw audio and video streams and replies in text. The authors build OmniVChat‑Studio, a multi‑agent simulator that generates single‑ and multi‑turn audio‑visual conversations, and use it to create OmniVChat‑Bench, a benchmark covering five dialogue abilities. They also propose OmniVChat‑RL, a reinforcement‑learning reward that b…

    Hugging Face Daily Papersarxiv.org2 minpaper