Hugging Face Daily PapersChia-Yuan Chang, Renyuan Cheng, Rui Feng1 min readpaperadvanced
Rufus-Air: An Open LLM Post-Training Recipe
Summary
Rufus‑Air presents a fully open, reproducible eight‑stage post‑training recipe for the 106‑billion‑parameter GLM‑4.5‑Air‑Base model, detailing data, reward design, and engineering choices. The pipeline—starting with diverse SFT and progressing through staged RL and RLHF—outperforms the official GLM‑4.5‑Air release and matches similarly sized open models.
- Eight‑stage pipeline (SFT → Reasoning RL → Coding RL → Instruction‑Following RL → General Agent → Coding Agent → Search Agent → RLHF) for GLM‑4.5‑Air‑Base (106B) outperforms the official release.
- Diverse, high‑quality SFT data creates a strong capability floor essential for later RL stages.
- Difficulty filtering of RL prompts keeps them in a productive learning range, reducing reward noise.
- Ordering stages by reward reliability—from hard verifiable rewards to softer judge‑based signals—improves convergence.
ML engineers building or fine‑tuning large language models will care because it offers a concrete, open‑source roadmap to achieve state‑of‑the‑art performance without proprietary data or teachers.
7/10


