proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersTianrun Yu, Kaixiang Zhao, Shangzhe Li1 min readpaperadvanced

Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

Summary

The authors model the training‑inference probability mismatch in RL‑fine‑tuned LLMs as an additive logit displacement and propose Calibrated Importance Sampling (CIS), which caps importance ratios using a confidence‑aware threshold. CIS provably bounds the variance term, adds a controllable bias, and achieves the highest average scores on several math‑reasoning benchmarks across three MoE models.

  • Mismatch between inference and training engines can be expressed as an additive logit displacement εₜ that is roughly independent of token confidence.
  • CIS applies a single positive‑displacement threshold, yielding importance‑ratio caps that tighten as token confidence grows.
  • Theoretical analysis shows CIS replaces the unbounded second‑moment term of exact importance sampling with a constant‑bounded term, with bias limited by the truncated excess.
  • Empirically, CIS outperforms baselines on five mathematical reasoning benchmarks across three mixture‑of‑experts LLMs, achieving the highest average score.

LLM engineers fine‑tuning with reinforcement learning need a low‑variance, bias‑controlled way to reconcile training‑inference probability gaps; CIS provides that.

8/10

Related reading

  1. Towards Full Pipeline FP8 Reinforcement Learning for LLMs

    FP8 quantization of RL pipelines for LLMs causes training instability because quantization noise skews importance ratios, zeroing gradients for negative‑advantage tokens and leading to entropy spikes. The authors introduce Calibrated Clipping, a dynamic adjustment of FP8 clipping bounds that matches BF16 quantile statistics, restoring stable training across GRPO and DAPO methods for 8‑32 B models…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Learning to solve hard problems in RL for LLMs by never giving up

    The post introduces the *Matthew Effect* in RL‑fine‑tuning of LLMs—performance gains concentrate on tasks the model already solves— and proposes *Never Give Up* (NGU), an adaptive sampling scheme that uses a small k for easy prompts and retries hard prompts with a high‑probability “never give up” loop. Experiments on math (AIME, GSM8k), code (Manufactoria), and larger‑scale setups (DeepScaler) sh…

    Hacker News front pagegithub.io11 minHN1179
  3. DataFlex-RL: An Evaluation Platform for RLVR Data Policies

    The paper introduces DataFlex‑RL, a platform to benchmark how different data‑selection policies affect reinforcement‑learning‑with‑verifiable‑rewards training. Across extensive experiments on Qwen2.5‑7B and Llama‑3.1‑8B, uniform sampling is the only method that consistently improves performance, and no alternative policy yields a statistically significant gain.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Sharpening Tax in Post-Training

    The authors introduce Sharpening Tax, a metric that measures how much RL post‑training reduces test‑time scalability (pass@K) of LLM agents. They also propose Posterior‑Tempered Group Sampling, a simple temperature‑adaptation technique that lowers this tax and improves both single‑shot accuracy and coverage.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

    PPO critics in reinforcement learning for LLMs suffer from "Value Flattening," where predicted state values are too flat compared to actual values. This paper identifies the causes as an implicit variance penalty and redundant updates, and proposes SParse Proximal Policy Optimization (SP3O) to mitigate it by supervising only a few well-separated states.

    Hugging Face Daily Papersarxiv.org1 minpaper