proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJinfa Huang, Jianming Xu, Jingyang Lin1 min readpaperadvanced

VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents

Summary

Long-form video agents often suffer from 'semantic thrashing' where growing append-only memory causes attention collapse and loss of key evidence. VideoLoop proposes a two-loop memory management system that retrieves artifacts from a filesystem and rewrites a bounded working memory to mitigate this issue.

  • Append-only working memory in long-form video agents leads to 'semantic thrashing' by accumulating noise and losing focus on key evidence.
  • VideoLoop employs an outer reasoning loop and an inner loop that retrieves past observations and rewrites a bounded working memory.
  • The method improves four popular LVLM backbones by an average of 4.2% points on the VideoMME (long) benchmark.
  • A blind judge found VideoLoop's managed context led to 81.1% correct answers on hard VideoMME questions, compared to 60.9% for append-only agents.

Engineers developing multimodal agents for long-form video understanding should care, as VideoLoop offers a concrete architectural solution to a fundamental memory management problem that hinders agent performance.

8/10

Related reading

  1. Automating coherent long-form video generation

    Google Research introduces a unified multi-agent framework to autonomously generate temporally consistent, long-form video narratives. It overcomes identity drift and cascading failures of current linear AI pipelines by treating generation as a global optimization and world-state tracking problem.

    Google Researchresearch.google10 min
  2. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

    APM-Bench is a new benchmark for evaluating persistent memory in egocentric streaming video assistants across intermittent sessions. It reveals a significant utility-latency-storage trade-off, showing current models struggle with long-term recall, low overhead, and proactive assistance simultaneously.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

    PoS is an inference-time framework that constructs and maintains explicit belief states for LLM agents, combining world state and unresolved task requirements. It validates consistency and detects "Belief Trapping" to ensure progress, achieving superior performance on long-horizon execution and diagnosis benchmarks across multiple LLM backbones.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. VideoGen-Agent: Reinforcing Video Generation Agents

    VideoGen-Agent is a multimodal RL‑trained agent that orchestrates augmentation, generation, and verification tools to improve text‑to‑video synthesis on a new 600‑prompt benchmark (VABench). It lifts a base generator’s score from 56.5 to 75.6 (‑19.1 pts) and to 86.1 when the generation tools are upgraded, with 84.3% human preference over the strongest baseline.

    Hugging Face Daily Papersarxiv.org1 minpaper