Hugging Face Daily PapersJinfa Huang, Jianming Xu, Jingyang Lin1 min readpaperadvanced
VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents
Summary
Long-form video agents often suffer from 'semantic thrashing' where growing append-only memory causes attention collapse and loss of key evidence. VideoLoop proposes a two-loop memory management system that retrieves artifacts from a filesystem and rewrites a bounded working memory to mitigate this issue.
- Append-only working memory in long-form video agents leads to 'semantic thrashing' by accumulating noise and losing focus on key evidence.
- VideoLoop employs an outer reasoning loop and an inner loop that retrieves past observations and rewrites a bounded working memory.
- The method improves four popular LVLM backbones by an average of 4.2% points on the VideoMME (long) benchmark.
- A blind judge found VideoLoop's managed context led to 81.1% correct answers on hard VideoMME questions, compared to 60.9% for append-only agents.
Engineers developing multimodal agents for long-form video understanding should care, as VideoLoop offers a concrete architectural solution to a fundamental memory management problem that hinders agent performance.
8/10
