Hugging Face Daily PapersZhiwei Zhang, Huayu Deng, Fei Zhao1 min readpaperadvanced
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
Summary
ROSS is a method for large language model post-training that selectively supervises historical self-generated rollouts, applying loss only to useful continuations while preserving full trajectory context. It consistently improves LLM performance across various tasks like code generation and instruction following, achieving gains on Qwen3.6-35B-A3B without requiring new policy rollouts.
- ROSS reuses historical self-generated rollouts from RL and on-policy distillation for LLM post-training.
- It applies loss only to selected, useful continuations within full historical trajectories, avoiding stale data.
- This method improves LLM performance without needing additional policy rollouts, saving computational resources.
- ROSS improved Qwen3.6-35B-A3B's MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%.
Researchers and practitioners in LLM fine-tuning should care, as ROSS offers a resource-efficient way to improve model performance by intelligently leveraging existing training data.
8/10