Hugging Face Daily PapersTianrun Yu, Kaixiang Zhao, Shangzhe Li1 min readpaperadvanced
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
Summary
The authors model the training‑inference probability mismatch in RL‑fine‑tuned LLMs as an additive logit displacement and propose Calibrated Importance Sampling (CIS), which caps importance ratios using a confidence‑aware threshold. CIS provably bounds the variance term, adds a controllable bias, and achieves the highest average scores on several math‑reasoning benchmarks across three MoE models.
- Mismatch between inference and training engines can be expressed as an additive logit displacement εₜ that is roughly independent of token confidence.
- CIS applies a single positive‑displacement threshold, yielding importance‑ratio caps that tighten as token confidence grows.
- Theoretical analysis shows CIS replaces the unbounded second‑moment term of exact importance sampling with a constant‑bounded term, with bias limited by the truncated excess.
- Empirically, CIS outperforms baselines on five mathematical reasoning benchmarks across three mixture‑of‑experts LLMs, achieving the highest average score.
LLM engineers fine‑tuning with reinforcement learning need a low‑variance, bias‑controlled way to reconcile training‑inference probability gaps; CIS provides that.
8/10