proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersWeihao Liu, Huangjie Zheng, Tianrong Chen1 min readpaperadvanced

Decoding Looped Transformers Better for (Almost) Free

Summary

LoopCD is a training‑free contrastive decoding technique for looped Transformers that leverages earlier recurrent states as weak predictions to guide token selection, either via an extra logit pass (LoopCD‑Logits) or directly in hidden‑state space (LoopCD‑Hidden). It improves pass@1 on code benchmarks by up to 9% absolute and enables halving recurrent loops, cutting inference FLOPs by 22‑48%.

  • LoopCD contrasts the final token prediction with an earlier loop's output, requiring no extra training.
  • LoopCD‑Logits adds one extra output pass per step; LoopCD‑Hidden uses hidden states with zero extra compute.
  • Achieves 61.88%→73.33% AIME 2024 pass@1 (Ouro‑2.6B) and 22.56%→31.71% HumanEval pass@1 (Huginn).
  • Halving the number of recurrent loops matches or exceeds full‑depth baselines, saving 22.5%‑48.2% FLOPs.

Engineers deploying looped Transformer models can boost generation quality and cut inference cost without retraining, directly improving production LLM services.

8/10

Related reading

  1. Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

    The paper demonstrates that transformer LLMs exhibit a linear superposition property where combined inputs produce a blended next‑token distribution, and that lightweight fine‑tuning can restore this linearity. It also introduces a guided decoding algorithm that extracts two distinct, coherent continuations from one forward pass.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus

    This paper investigates the "lossless" claim of Orthrus, a hybrid architecture for accelerating LLM inference. It finds that under BF16 precision, Orthrus diverges from the exact autoregressive output trajectory in over 50% of cases, though FP32 maintains exact matching. Despite BF16 divergence, downstream task performance was not systematically degraded.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Scaling Laws for Looped Mixture of Experts

    This paper introduces Loop Scaling Laws, which jointly model recurrence and sparsity in Mixture-of-Experts (MoE) transformers. It finds that recurrence and sparsity offer complementary efficiency gains, improving prediction accuracy and enabling more efficient large model design.

    Hugging Face Daily Papersarxiv.org1 minpaperHN2
  4. Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

    Decoy Direction Optimization (DDO) is a post‑hoc weight‑editing defense for open‑weight LLMs that injects a high‑magnitude nonlinear decoy into MLP neurons, corrupting contrastive estimators used by Refusal Feature Ablation (RFA) attacks. The paper proves a spectral bound on the effect, evaluates DDO on six model families (including Llama‑3‑8B‑Instruct), and shows <10 % attack success rate (ASR)…

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

    WhiteMatter introduces all-to-all cross-layer connections in Transformers by mixing past-token representations from any depth into shared KV cache channels. This approach allows for performance comparable to 50% larger standard Transformers or improved performance with half the KV cache size. It also addresses training slowdowns with a novel cyclic iteration method.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

    Flash-dLLM is a training-free framework that accelerates Diffusion LLM inference by addressing GPU memory I/O bottlenecks with an I/O-aware KV-cache kernel. It also introduces a KV-cache-driven draft-and-verify decoding strategy, achieving significant speedups (up to 11x) over prior methods.

    Hugging Face Daily Papersarxiv.org1 minpaper