proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersNagham Omar, Mahmoud Jabarin, Maya Rozenshtein1 min readpaperadvanced

Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs

Summary

This paper redefines LLM generalization as output stability across varied inputs, rather than just accuracy. It introduces SAGO, a multi-axis framework, finding that many LLMs exhibit significant and consistent generalization instability across different behavioral dimensions.

  • Generalization in LLMs should be defined as output stability across varied inputs, not just aggregate accuracy.
  • Existing evaluation methods conflate robustness with overall benchmark performance, obscuring true generalization.
  • The SAGO framework measures variability across generation consistency, internal activations, confidence, and response mirroring.
  • Many commonly used LLMs exhibit statistically significant and consistent generalization instability.

Engineers and researchers evaluating or deploying LLMs should care, as it provides a more robust and nuanced method for assessing model generalization beyond simple accuracy metrics.

8/10

Related reading

  1. Learning to solve hard problems in RL for LLMs by never giving up

    The post introduces the *Matthew Effect* in RL‑fine‑tuning of LLMs—performance gains concentrate on tasks the model already solves— and proposes *Never Give Up* (NGU), an adaptive sampling scheme that uses a small k for easy prompts and retries hard prompts with a high‑probability “never give up” loop. Experiments on math (AIME, GSM8k), code (Manufactoria), and larger‑scale setups (DeepScaler) sh…

    Hacker News front pagegithub.io11 minHN1179
  2. Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models

    The paper reframes fine‑tuning of instruction‑tuned LLMs as a direction‑selection problem under a fixed behavioral‑drift budget, showing that the update direction, not magnitude, determines trade‑offs between target performance and capability preservation. In QA‑only fine‑tuning of Qwen‑3 models, layer‑selective probing finds effective directions that boost scientific reasoning and multilingual t…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

    The paper demonstrates that transformer LLMs exhibit a linear superposition property where combined inputs produce a blended next‑token distribution, and that lightweight fine‑tuning can restore this linearity. It also introduces a guided decoding algorithm that extracts two distinct, coherent continuations from one forward pass.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. How Value Induction Reshapes LLM Behaviour

    Apple researchers fine‑tune LLMs on curated subsets of value‑oriented preference data and measure cross‑value effects, safety, and anthropomorphic language. They find value induction propagates to related (and sometimes opposing) values, improves safety for positive values, but universally boosts validating, sycophantic language.

    Apple Machine Learning Researchapple.com1 minpaper
  5. Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

    Increasing the number of LLM candidates (N) improves reasoning accuracy, but the way those candidates are generated (batch size vs sequential calls) dramatically affects latency, GPU‑hours, and energy. On A100 GPUs, eight serial 1‑candidate calls consume ~5× more energy and take ~6× longer than a single 8‑candidate batched call, even though total candidate count is identical. The authors recommen…

    Hugging Face Daily Papersarxiv.org2 minpaper