proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersZhe Feng, Longfei Liu, Wei Liu1 min readpaperadvanced

GTR: Gated Token Recurrence for Efficient Dense Prediction

Summary

GTR replaces quadratic softmax attention with gated linear attention and spatial recurrence, delivering a fast, softmax‑free vision backbone. It reaches 58.9 COCO box AP with ~1.9 ms latency on RTX 4090 and runs efficiently on edge GPUs via TensorRT.

  • Gated Token Recurrence (GTR) uses gated linear attention and alternating spatial scans to avoid the O(N²) cost of global softmax attention.
  • Training relies only on final‑layer patch‑token alignment with a linear projection and L2 loss, eliminating masked‑token or intermediate supervision.
  • GTR‑L hits 58.9 box AP on COCO val2017 while achieving 1.908 ms median batch‑one latency on compiled FP16 RTX 4090.
  • A custom chunkwise CUDA kernel is 4× faster than FLA v0.5.0 at 1.6K tokens, and TensorRT deployment on DRIVE AGX Thor yields 2.3‑8.8 ms latency across tasks.

Vision engineers needing high‑resolution dense prediction at real‑time speeds, especially for edge or GPU‑constrained deployments, should see GTR as a practical alternative to softmax attention.

7/10

Related reading

  1. VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

    VC-Attention introduces a training‑free low‑bit attention pipeline for diffusion transformers. It smooths value tensors via lightweight online clustering (V‑Smooth) and quantizes only the residual after subtracting block means, restoring the mean from the softmax row sum. It also replaces the FP32 softmax exponential with a fused FP8 cast (ExpCast‑FP8) that maps log‑scores directly to E4M3 probab…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs

    This post details how Google optimized spatio-temporal attention for video diffusion models on TPUs, turning theoretical sparsity into actual inference speedups. Key optimizations include specializing tile execution paths and carefully tuning tile sizes, resulting in significant latency reductions compared to dense attention.

    Google Developersgoogleblog.com11 min
  3. Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

    Video DeltaNet (VDN) replaces full‑softmax attention in video diffusion models with a hybrid: per‑frame local Softmax for fine detail and a bidirectional linear memory (Video Delta Attention) for long‑range context. A teacher‑alignment schedule injects the linear branch into a pretrained MiniMax H3 model, preserving Softmax for text/audio streams. On eight NVIDIA B200 GPUs VDN‑H3 denoises a 14.3‑…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

    This paper evaluates GPT-6 Astra and five other frontier general-purpose AI systems across 34 computer vision capabilities and 55 benchmarks. It finds these systems excel at semantic interpretation and reasoning, but struggle with metric geometric accuracy, faithful reconstruction, and fine-grained specialized knowledge.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

    VisionHOPE introduces a visual backbone that updates its own parameters on‑the‑fly using five coupled memories, with a stability‑matched step‑size scheme that keeps updates non‑expansive. The approach attains competitive ImageNet, COCO and ADE20K performance, showing self‑modifying learning systems can serve as practical general‑purpose vision models.

    Hugging Face Daily Papersarxiv.org1 minpaper