Hugging Face Daily PapersJuyi Lin, Zhiqiang Lao, Jiali Cui1 min readpaperadvanced
LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception
Summary
LEAP is a framework for long audio-video question answering that addresses context limits by retrieving relevant evidence blocks. It uses a two-stage process of block-wise localization and a bounded answer pass, achieving significant performance improvements on AVQA benchmarks.
- LEAP decouples evidence localization from reasoning for hour-scale audio-video inputs.
- It divides recordings into fixed blocks, applying a lightweight pass to score and select short candidate windows.
- Only the highest-ranked windows are re-encoded for the final answer pass, keeping context size independent of recording duration.
- Localization can use pre-computed transcripts, while the final answer pass uses raw audio-visual streams for fine-grained evidence.
Engineers developing multimodal AI systems for long-form audio-video content will find this relevant for improving efficiency and accuracy in question answering by managing context effectively.
8/10