1
GTR: Gated Token Recurrence for Efficient Dense Prediction
GTR replaces quadratic softmax attention with gated linear attention and spatial recurrence, delivering a fast, softmax‑free vision backbone. It reaches 58.9 COCO box AP with ~1.9 ms latency on RTX 4090 and runs efficiently on edge GPUs via TensorRT.
Hugging Face Daily Papersarxiv.org1 minpaper
