proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersAyush Jain, Sreeharsha Paruchuri, Ishita Gupta1 min readpaperadvanced

TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

Summary

TrackEverything is a novel 3D point tracker that overcomes the trade-off between sparse long-horizon tracking and dense short-clip tracking. It represents videos as persistent 3D scene tracks, scaling with unique physical geometry rather than video duration, and achieves dense tracking for over 1000 frames within 40 GB of GPU memory.

  • Represents videos as persistent 3D scene tracks in world coordinates, decoupling complexity from video duration.
  • Employs voxelization-based de-duplication at sliding-window boundaries to merge co-located tracks.
  • Decomposes tracking into an endpoint refiner (destination, static/dynamic classification) and a lightweight trajectory refiner for dynamic points.
  • Introduces 3D WAFT, replacing memory-intensive 4D correlation volumes with efficient feature sampling in the scene cloud.

This paper is significant for researchers and engineers working on computer vision and robotics, as it enables dense, long-duration 3D tracking with practical memory constraints, opening new possibilities for applications like autonomous systems and augmented reality.

8/10

Related reading

  1. Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

    The paper investigates why video diffusion models often break physical laws, pinpointing excessive spatial decay from Rotary Position Embedding (RoPE) as the culprit. By analyzing cross‑attention trajectories and self‑attention patterns, the authors identify specific attention heads that drive motion planning. They propose a lightweight fix: scaling RoPE frequency per denoising step, which empiri…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation

    GEAR is a geometry‑enabled attention routing framework for long‑horizon camera‑controlled video generation. It treats per‑frame geometry as token‑level addresses, uses Geometric Correspondence Attention and an Invisible Octree to retrieve visual memory, achieving state‑of‑the‑art quality and consistent control over minute‑long videos.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

    Video DeltaNet (VDN) replaces full‑softmax attention in video diffusion models with a hybrid: per‑frame local Softmax for fine detail and a bidirectional linear memory (Video Delta Attention) for long‑range context. A teacher‑alignment schedule injects the linear branch into a pretrained MiniMax H3 model, preserving Softmax for text/audio streams. On eight NVIDIA B200 GPUs VDN‑H3 denoises a 14.3‑…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

    APM-Bench is a new benchmark for evaluating persistent memory in egocentric streaming video assistants across intermittent sessions. It reveals a significant utility-latency-storage trade-off, showing current models struggle with long-term recall, low overhead, and proactive assistance simultaneously.

    Hugging Face Daily Papersarxiv.org1 minpaper