Hugging Face Daily PapersChangdae Oh, Qi Zeng, Qi Qi1 min readpaperadvanced
Sharpening Tax in Post-Training
Summary
The authors introduce Sharpening Tax, a metric that measures how much RL post‑training reduces test‑time scalability (pass@K) of LLM agents. They also propose Posterior‑Tempered Group Sampling, a simple temperature‑adaptation technique that lowers this tax and improves both single‑shot accuracy and coverage.
- Sharpening Tax quantifies the loss in pass@K scalability after RL post‑training and can be estimated from a few rollouts.
- Pre‑trained LLMs with a lightweight inference harness often achieve higher solution coverage than post‑trained models given sufficient sampling budget.
- RL post‑training pushes task performance to a bimodal distribution, improving consistency but hurting coverage.
- Posterior‑Tempered Group Sampling adapts temperature per prompt based on estimated difficulty, reducing Sharpening Tax.
LLM engineers and RL researchers should care because the paper reveals a coverage trade‑off in post‑training and offers a cheap sampling tweak to recover it.
8/10