proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJi Xia, Tingting Liao, Xuezhi Liang1 min readpaperadvanced

LOCI: Spatial Linear Memory for Streaming World Models

Summary

LOCI is a hybrid spatial-memory architecture for streaming video world models, combining key-value caches with recurrent linear attention. It leverages projective camera geometry to condition memory operations, enabling more faithful reproduction of revisited content and significantly reducing peak memory usage for long videos.

  • LOCI employs a hybrid memory: key-value caches for visual detail and recurrent linear attention for compact history.
  • Memory reads and writes are conditioned on projective camera geometry, integrating viewpoint into memory addressing and content.
  • The architecture reproduces revisited content more faithfully than representative world models and full-softmax baselines.
  • It lowers peak memory by approximately 30% for full history and supports constant memory streaming with bounded observations.

Engineers and researchers working on video world models for applications like robotics or AR/VR should care about LOCI's novel approach to managing long-term memory and visual fidelity efficiently.

8/10

Related reading

  1. Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation

    GEAR is a geometry‑enabled attention routing framework for long‑horizon camera‑controlled video generation. It treats per‑frame geometry as token‑level addresses, uses Geometric Correspondence Attention and an Invisible Octree to retrieve visual memory, achieving state‑of‑the‑art quality and consistent control over minute‑long videos.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

    APM-Bench is a new benchmark for evaluating persistent memory in egocentric streaming video assistants across intermittent sessions. It reveals a significant utility-latency-storage trade-off, showing current models struggle with long-term recall, low overhead, and proactive assistance simultaneously.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

    Video DeltaNet (VDN) replaces full‑softmax attention in video diffusion models with a hybrid: per‑frame local Softmax for fine detail and a bidirectional linear memory (Video Delta Attention) for long‑range context. A teacher‑alignment schedule injects the linear branch into a pretrained MiniMax H3 model, preserving Softmax for text/audio streams. On eight NVIDIA B200 GPUs VDN‑H3 denoises a 14.3‑…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

    The paper presents SemanTok, a flexible video tokenizer that injects frozen DINO features and reconstructs them from any token prefix, achieving strong semantic alignment and fidelity. A 201 M SemanTok AR model matches or exceeds a 3.4× larger VideoFlexTok baseline, with cheaper short‑prefix prediction and better generation quality.

    Hugging Face Daily Papersarxiv.org1 minpaper