Hugging Face Daily PapersJi Xia, Tingting Liao, Xuezhi Liang1 min readpaperadvanced
LOCI: Spatial Linear Memory for Streaming World Models
Summary
LOCI is a hybrid spatial-memory architecture for streaming video world models, combining key-value caches with recurrent linear attention. It leverages projective camera geometry to condition memory operations, enabling more faithful reproduction of revisited content and significantly reducing peak memory usage for long videos.
- LOCI employs a hybrid memory: key-value caches for visual detail and recurrent linear attention for compact history.
- Memory reads and writes are conditioned on projective camera geometry, integrating viewpoint into memory addressing and content.
- The architecture reproduces revisited content more faithfully than representative world models and full-softmax baselines.
- It lowers peak memory by approximately 30% for full history and supports constant memory streaming with bounded observations.
Engineers and researchers working on video world models for applications like robotics or AR/VR should care about LOCI's novel approach to managing long-term memory and visual fidelity efficiently.
8/10