proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersChangdae Oh, Qi Zeng, Qi Qi1 min readpaperadvanced

Sharpening Tax in Post-Training

Summary

The authors introduce Sharpening Tax, a metric that measures how much RL post‑training reduces test‑time scalability (pass@K) of LLM agents. They also propose Posterior‑Tempered Group Sampling, a simple temperature‑adaptation technique that lowers this tax and improves both single‑shot accuracy and coverage.

  • Sharpening Tax quantifies the loss in pass@K scalability after RL post‑training and can be estimated from a few rollouts.
  • Pre‑trained LLMs with a lightweight inference harness often achieve higher solution coverage than post‑trained models given sufficient sampling budget.
  • RL post‑training pushes task performance to a bimodal distribution, improving consistency but hurting coverage.
  • Posterior‑Tempered Group Sampling adapts temperature per prompt based on estimated difficulty, reducing Sharpening Tax.

LLM engineers and RL researchers should care because the paper reveals a coverage trade‑off in post‑training and offers a cheap sampling tweak to recover it.

8/10

Related reading

  1. Mitigating the Length-Scaling Tax with Online Distillation

    The authors define the length‑scaling tax (LST) as excess response length without accuracy gain and propose Length Self‑Distillation (LSD), an online EMA‑based teacher that requires no external model. Experiments show LSD matches or exceeds RL performance while cutting LST from 19% to -3.7% on single‑turn and from 31.4% to 13.7% on multi‑turn tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Learning to solve hard problems in RL for LLMs by never giving up

    The post introduces the *Matthew Effect* in RL‑fine‑tuning of LLMs—performance gains concentrate on tasks the model already solves— and proposes *Never Give Up* (NGU), an adaptive sampling scheme that uses a small k for easy prompts and retries hard prompts with a high‑probability “never give up” loop. Experiments on math (AIME, GSM8k), code (Manufactoria), and larger‑scale setups (DeepScaler) sh…

    Hacker News front pagegithub.io11 minHN1179
  3. AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

    AutoDataBench introduces a new benchmark to evaluate if AI agents can autonomously generate high-quality training data for LLMs, judged against production-like acceptance criteria. The study found that current agents can produce usable tasks but struggle significantly with efficiency, scoring low on a time-constrained budget.

    Hugging Face Daily Papersarxiv.org2 minpaper
  4. ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

    ROSS is a method for large language model post-training that selectively supervises historical self-generated rollouts, applying loss only to useful continuations while preserving full trajectory context. It consistently improves LLM performance across various tasks like code generation and instruction following, achieving gains on Qwen3.6-35B-A3B without requiring new policy rollouts.

    Hugging Face Daily Papersarxiv.org1 minpaper