proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersZesong Yang, Weikai Chen, Liyuan Cui1 min readpaperadvanced

Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation

Summary

GEAR is a geometry‑enabled attention routing framework for long‑horizon camera‑controlled video generation. It treats per‑frame geometry as token‑level addresses, uses Geometric Correspondence Attention and an Invisible Octree to retrieve visual memory, achieving state‑of‑the‑art quality and consistent control over minute‑long videos.

  • GEAR uses per‑frame geometry as token‑level addresses, avoiding costly global 3D fusion and error accumulation.
  • Geometric Correspondence Attention injects matched historical features into noisy target patches during diffusion denoising.
  • An Invisible Octree accumulates visibility evidence to reject occluded correspondences.
  • The method attains state‑of‑the‑art visual quality, precise camera control, and revisit consistency on minute‑long videos.

Engineers building diffusion‑based video generators or long‑trajectory camera control need efficient, accurate memory access, which GEAR provides.

8/10

Related reading

  1. LOCI: Spatial Linear Memory for Streaming World Models

    LOCI is a hybrid spatial-memory architecture for streaming video world models, combining key-value caches with recurrent linear attention. It leverages projective camera geometry to condition memory operations, enabling more faithful reproduction of revisited content and significantly reducing peak memory usage for long videos.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

    Video DeltaNet (VDN) replaces full‑softmax attention in video diffusion models with a hybrid: per‑frame local Softmax for fine detail and a bidirectional linear memory (Video Delta Attention) for long‑range context. A teacher‑alignment schedule injects the linear branch into a pretrained MiniMax H3 model, preserving Softmax for text/audio streams. On eight NVIDIA B200 GPUs VDN‑H3 denoises a 14.3‑…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Octrees as an Explicit 3D Language

    OctLLM treats 3D geometry as a sequence of octree occupancy tokens, using a Sparse Octree to keep sequences short while preserving shape. It adds lightweight 3D branches to a frozen vision‑language backbone, achieving state‑of‑the‑art image‑to‑3D generation with far fewer trainable parameters.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

    TrackEverything is a novel 3D point tracker that overcomes the trade-off between sparse long-horizon tracking and dense short-clip tracking. It represents videos as persistent 3D scene tracks, scaling with unique physical geometry rather than video duration, and achieves dense tracking for over 1000 frames within 40 GB of GPU memory.

    Hugging Face Daily Papersarxiv.org1 minpaper