proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersPavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk1 min readpaperadvanced

Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

Summary

The paper demonstrates that transformer LLMs exhibit a linear superposition property where combined inputs produce a blended next‑token distribution, and that lightweight fine‑tuning can restore this linearity. It also introduces a guided decoding algorithm that extracts two distinct, coherent continuations from one forward pass.

  • Linear superposition: linearly mixing two input streams yields a next‑token distribution close to the average of the individual distributions.
  • Superposition is intrinsic to the Transformer architecture but diminishes as pretraining progresses.
  • Lightweight fine‑tuning restores linearity, significantly reducing divergence between superposed and averaged outputs.
  • Guided decoding can disentangle the superposed output, producing two coherent continuations from a single forward pass.

LLM engineers and researchers should care because the technique enables efficient multi‑output generation and offers insight into transformer internals for better decoding strategies.

7/10

Related reading

  1. Attention Is All You Need

    The paper proposes the Transformer, a sequence‑to‑sequence model that relies solely on self‑attention, eliminating recurrence and convolutions. It achieves state‑of‑the‑art translation BLEU scores while training orders of magnitude faster.

    Hall of Famearxiv.org27 minpaper
  2. Transformers Explained Visually

    The article walks through the core components of a text‑generative Transformer—embedding, multi‑head self‑attention, MLP, and output projection—using GPT‑2 small as a concrete example. It shows the dimensions, parameter counts, and step‑by‑step calculations that underlie token prediction.

    Hacker News front pagegithub.io11 minHN63388
  3. WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

    WhiteMatter introduces all-to-all cross-layer connections in Transformers by mixing past-token representations from any depth into shared KV cache channels. This approach allows for performance comparable to 50% larger standard Transformers or improved performance with half the KV cache size. It also addresses training slowdowns with a novel cyclic iteration method.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. A theoretical separation between quantum computers & LLMs

    The IBM research blog explains two new theoretical results that prove shallow constant‑depth quantum circuits can outperform decoder‑only transformers on a functional task (iterated index) and diffusion language models on a sampling task (parity‑sampling). The proofs give asymptotic separations but are not yet practical.

    IBM Researchibm.com6 min