Hugging Face Daily PapersZhe Feng, Longfei Liu, Wei Liu1 min readpaperadvanced
GTR: Gated Token Recurrence for Efficient Dense Prediction
Summary
GTR replaces quadratic softmax attention with gated linear attention and spatial recurrence, delivering a fast, softmax‑free vision backbone. It reaches 58.9 COCO box AP with ~1.9 ms latency on RTX 4090 and runs efficiently on edge GPUs via TensorRT.
- Gated Token Recurrence (GTR) uses gated linear attention and alternating spatial scans to avoid the O(N²) cost of global softmax attention.
- Training relies only on final‑layer patch‑token alignment with a linear projection and L2 loss, eliminating masked‑token or intermediate supervision.
- GTR‑L hits 58.9 box AP on COCO val2017 while achieving 1.908 ms median batch‑one latency on compiled FP16 RTX 4090.
- A custom chunkwise CUDA kernel is 4× faster than FLA v0.5.0 at 1.6K tokens, and TensorRT deployment on DRIVE AGX Thor yields 2.3‑8.8 ms latency across tasks.
Vision engineers needing high‑resolution dense prediction at real‑time speeds, especially for edge or GPU‑constrained deployments, should see GTR as a practical alternative to softmax attention.
7/10
