1
The information geometry of large language models is shared, learned, and controllable
The paper shows that the Fisher‑Rao geometry of next‑token probabilities is largely shared across transformer, state‑space and recurrent LLMs, and that this shared geometry can be used to design low‑disturbance interventions that steer model behavior. Experiments demonstrate that geometry predicts semantic transfer, fact acquisition, and enables reusable control better than Euclidean methods.
Hugging Face Daily Papersarxiv.org1 minpaper
