Hugging Face Daily PapersPavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk1 min readpaperadvanced
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
Summary
The paper demonstrates that transformer LLMs exhibit a linear superposition property where combined inputs produce a blended next‑token distribution, and that lightweight fine‑tuning can restore this linearity. It also introduces a guided decoding algorithm that extracts two distinct, coherent continuations from one forward pass.
- Linear superposition: linearly mixing two input streams yields a next‑token distribution close to the average of the individual distributions.
- Superposition is intrinsic to the Transformer architecture but diminishes as pretraining progresses.
- Lightweight fine‑tuning restores linearity, significantly reducing divergence between superposed and averaged outputs.
- Guided decoding can disentangle the superposed output, producing two coherent continuations from a single forward pass.
LLM engineers and researchers should care because the technique enables efficient multi‑output generation and offers insight into transformer internals for better decoding strategies.
7/10

