Hugging Face Daily PapersZehao Jin, Ruixuan Deng, Junran Wang1 min readpaperadvanced
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Summary
The paper shows that standard pretrained transformers stop using deeper layers after only a few lines of context. Adding a rank‑8 LoRA to an early layer while keeping the rest frozen dramatically extends effective reasoning depth, boosting exact accuracy on long‑chain tasks from ~15% to near‑100%.
- Pretrained models reliably follow only 1.4–3.6 lines of context, even after extra pretraining loops.
- A rank‑8 LoRA added to an early layer (all other weights frozen) extends usable depth to >15 lines on Qwen3‑8B, achieving 99% exact accuracy.
- The LoRA creates a relay where middle layers propagate chain identity, letting frozen heads read farther up the chain.
- Removing parent‑line attention breaks the relay, confirming its role in extended reasoning.
LLM engineers and researchers should care because a tiny, cheap LoRA edit can unlock far deeper reasoning without full model retraining.
7/10