proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersXuanyi Zhou, Qiuyang Mang, Huanzhi Mao1 min readpaperadvanced

EasyPPO: Stabilizing the Critic Is Key

Summary

EasyPPO addresses two key instability issues in PPO's critic when training large language models: biased filtering of truncated rollouts and heterogeneous return noise. It introduces actor-only overlong filtering, noise-normalized critic regression, and smaller critic mini-batches, achieving significant performance gains over vanilla PPO across various tasks.

  • PPO's learned critic, while reducing variance, is a major source of instability when training large language models.
  • Two critic failure modes are identified: biased filtering of truncated rollouts and heterogeneous return noise in finite batches.
  • EasyPPO uses actor-only overlong filtering to train the critic on returns from both completed and truncated rollouts.
  • Noise-normalized critic regression weights each prompt's critic loss by the inverse standard deviation of its sampled returns.

Engineers and researchers working on reinforcement learning for large language models should care, as EasyPPO offers a more stable and performant PPO training approach.

8/10

Related reading

  1. Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

    PPO critics in reinforcement learning for LLMs suffer from "Value Flattening," where predicted state values are too flat compared to actual values. This paper identifies the causes as an implicit variance penalty and redundant updates, and proposes SParse Proximal Policy Optimization (SP3O) to mitigate it by supervising only a few well-separated states.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. SteerDuplex: Steerable Duplex Speech Dialogue Models

    The paper presents SteerDuplex, a full‑duplex speech dialogue model that can be steered along tone, persona, and speed via instruction following, and introduces the SteerBench benchmark to evaluate such steerability. Supervised training yields a 44.5 % pass‑rate lift, and reinforcement‑learning fine‑tuning improves interruption handling and reduces pause barge‑ins, though reward hacking remains a…

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

    PLC‑DPO extends Direct Preference Optimization by using the policy‑reference margin to route each training pair into clean, flipped, or tie categories, actively correcting noisy or ambiguous labels. Across extensive benchmarks it improves mean win‑rate from 55.5 % to 60.5 % and stays stable under injected noise and tie stress tests.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

    DACA‑GRPO adds denoising‑aware credit assignment to GRPO‑style RL trainers for diffusion LLMs. It computes per‑token importance scores from intermediate denoising steps and uses stratified masking to reduce mean‑field bias in likelihood estimates. Plug‑and‑play on three existing GRPO methods, it yields consistent gains on seven downstream tasks (up to +5.6 pp math, +7.4 pp code, +36.3 pp constrai…

    Apple Machine Learning Researchapple.com1 minpaper
  5. Language Models are Few-Shot Learners

    This paper introduces GPT-3, a 175-billion-parameter autoregressive language model, demonstrating that scaling model size significantly improves few-shot learning. It achieves strong performance on many NLP tasks by conditioning on text instructions and examples, often without needing gradient updates or fine-tuning.

    Hall of Famearxiv.org166 minpaper