proomt

Search

Search posts, papers, and topics

training inference mismatch

RSS
  1. 1

    Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

    The authors model the training‑inference probability mismatch in RL‑fine‑tuned LLMs as an additive logit displacement and propose Calibrated Importance Sampling (CIS), which caps importance ratios using a confidence‑aware threshold. CIS provably bounds the variance term, adds a controllable bias, and achieves the highest average scores on several math‑reasoning benchmarks across three MoE models.

    Hugging Face Daily Papersarxiv.org1 minpaper