proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersZehao Jin, Ruixuan Deng, Junran Wang1 min readpaperadvanced

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Summary

The paper shows that standard pretrained transformers stop using deeper layers after only a few lines of context. Adding a rank‑8 LoRA to an early layer while keeping the rest frozen dramatically extends effective reasoning depth, boosting exact accuracy on long‑chain tasks from ~15% to near‑100%.

  • Pretrained models reliably follow only 1.4–3.6 lines of context, even after extra pretraining loops.
  • A rank‑8 LoRA added to an early layer (all other weights frozen) extends usable depth to >15 lines on Qwen3‑8B, achieving 99% exact accuracy.
  • The LoRA creates a relay where middle layers propagate chain identity, letting frozen heads read farther up the chain.
  • Removing parent‑line attention breaks the relay, confirming its role in extended reasoning.

LLM engineers and researchers should care because a tiny, cheap LoRA edit can unlock far deeper reasoning without full model retraining.

7/10

Related reading

  1. Attention Is All You Need

    The paper proposes the Transformer, a sequence‑to‑sequence model that relies solely on self‑attention, eliminating recurrence and convolutions. It achieves state‑of‑the‑art translation BLEU scores while training orders of magnitude faster.

    Hall of Famearxiv.org27 minpaper
  2. Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

    Latent Interface Training (LIT) first learns a goal‑conditioned action prior without visual input, then adds a pose‑supervised latent interface as the only visual conditioning path. Applied to several vision‑language‑action models, LIT cuts vision‑action shortcuts and lifts LIBERO‑Plus success by 3.9–10.7 points and real‑world task success by 13.3–16.7 points under distribution shifts.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

    The paper demonstrates that transformer LLMs exhibit a linear superposition property where combined inputs produce a blended next‑token distribution, and that lightweight fine‑tuning can restore this linearity. It also introduces a guided decoding algorithm that extracts two distinct, coherent continuations from one forward pass.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Transformers now runs llama.cpp quants

    Transformers now supports GGUF quantized checkpoints (used by llama.cpp) via the `kernels` library. Load a GGUF model with `from_pretrained` on Apple Silicon, run generation with standard Transformers APIs, or serve it via `transformers serve`. Benchmarks on an M2 Max show token‑throughput comparable to llama.cpp, with fallback to dequantization if kernels are missing. The integration also enable…

    Hugging Facehuggingface.co10 minHN4