Hugging Face Daily PapersWeihao Liu, Huangjie Zheng, Tianrong Chen1 min readpaperadvanced
Decoding Looped Transformers Better for (Almost) Free
Summary
LoopCD is a training‑free contrastive decoding technique for looped Transformers that leverages earlier recurrent states as weak predictions to guide token selection, either via an extra logit pass (LoopCD‑Logits) or directly in hidden‑state space (LoopCD‑Hidden). It improves pass@1 on code benchmarks by up to 9% absolute and enables halving recurrent loops, cutting inference FLOPs by 22‑48%.
- LoopCD contrasts the final token prediction with an earlier loop's output, requiring no extra training.
- LoopCD‑Logits adds one extra output pass per step; LoopCD‑Hidden uses hidden states with zero extra compute.
- Achieves 61.88%→73.33% AIME 2024 pass@1 (Ouro‑2.6B) and 22.56%→31.71% HumanEval pass@1 (Huginn).
- Halving the number of recurrent loops matches or exceeds full‑depth baselines, saving 22.5%‑48.2% FLOPs.
Engineers deploying looped Transformer models can boost generation quality and cut inference cost without retraining, directly improving production LLM services.
8/10