proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersZhiwei Zhang, Huayu Deng, Fei Zhao1 min readpaperadvanced

ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

Summary

ROSS is a method for large language model post-training that selectively supervises historical self-generated rollouts, applying loss only to useful continuations while preserving full trajectory context. It consistently improves LLM performance across various tasks like code generation and instruction following, achieving gains on Qwen3.6-35B-A3B without requiring new policy rollouts.

  • ROSS reuses historical self-generated rollouts from RL and on-policy distillation for LLM post-training.
  • It applies loss only to selected, useful continuations within full historical trajectories, avoiding stale data.
  • This method improves LLM performance without needing additional policy rollouts, saving computational resources.
  • ROSS improved Qwen3.6-35B-A3B's MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%.

Researchers and practitioners in LLM fine-tuning should care, as ROSS offers a resource-efficient way to improve model performance by intelligently leveraging existing training data.

8/10

Related reading

  1. Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

    This paper introduces Recursive Self-Rewrite (RSR), a framework that enables a single base LLM to discover solutions for complex tasks using diverse specialized environments (harnesses). It then reconstructs these successful trajectories into training data suitable for a general environment, significantly improving the model's performance on various benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Sharpening Tax in Post-Training

    The authors introduce Sharpening Tax, a metric that measures how much RL post‑training reduces test‑time scalability (pass@K) of LLM agents. They also propose Posterior‑Tempered Group Sampling, a simple temperature‑adaptation technique that lowers this tax and improves both single‑shot accuracy and coverage.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. EasyPPO: Stabilizing the Critic Is Key

    EasyPPO addresses two key instability issues in PPO's critic when training large language models: biased filtering of truncated rollouts and heterogeneous return noise. It introduces actor-only overlong filtering, noise-normalized critic regression, and smaller critic mini-batches, achieving significant performance gains over vanilla PPO across various tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models

    The paper reframes fine‑tuning of instruction‑tuned LLMs as a direction‑selection problem under a fixed behavioral‑drift budget, showing that the update direction, not magnitude, determines trade‑offs between target performance and capability preservation. In QA‑only fine‑tuning of Qwen‑3 models, layer‑selective probing finds effective directions that boost scientific reasoning and multilingual t…

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

    The paper systematically evaluates on‑policy vs off‑policy rollouts in LLM distillation, finding that rollout policy matters less than the KL direction and learning rate. Forward KL is robust to rollout choice, while reverse KL prefers student rollouts, and learning rate drives forgetting and sparsity.

    Hugging Face Daily Papersarxiv.org1 minpaper