proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersChia-Yuan Chang, Renyuan Cheng, Rui Feng1 min readpaperadvanced

Rufus-Air: An Open LLM Post-Training Recipe

Summary

Rufus‑Air presents a fully open, reproducible eight‑stage post‑training recipe for the 106‑billion‑parameter GLM‑4.5‑Air‑Base model, detailing data, reward design, and engineering choices. The pipeline—starting with diverse SFT and progressing through staged RL and RLHF—outperforms the official GLM‑4.5‑Air release and matches similarly sized open models.

  • Eight‑stage pipeline (SFT → Reasoning RL → Coding RL → Instruction‑Following RL → General Agent → Coding Agent → Search Agent → RLHF) for GLM‑4.5‑Air‑Base (106B) outperforms the official release.
  • Diverse, high‑quality SFT data creates a strong capability floor essential for later RL stages.
  • Difficulty filtering of RL prompts keeps them in a productive learning range, reducing reward noise.
  • Ordering stages by reward reliability—from hard verifiable rewards to softer judge‑based signals—improves convergence.

ML engineers building or fine‑tuning large language models will care because it offers a concrete, open‑source roadmap to achieve state‑of‑the‑art performance without proprietary data or teachers.

7/10

Related reading

  1. Saving Jet Fuel

    A step‑by‑step tutorial showing how to use the open‑source scikit‑decide framework together with the OpenAP aircraft performance model to compute fuel‑optimal flight trajectories. The post details the author’s high‑end workstation, installs Python 3.12, scikit‑decide, OpenAP, and DuckDB with several extensions, then explores OpenAP’s aircraft data (e.g., A380‑800 specs and drag polar) and demonst…

    Hacker News front pagemarksblogg.com26 minHN13876
  2. GLM 5.3 FlashX now available on AI Gateway

    Vercel AI Gateway now offers the GLM‑5.3‑FlashX model, a fast (~200 tps) multimodal coding LLM. The post includes a TypeScript streaming example, notes the model’s fit for coding agents and interactive tools, and lists AI Gateway features (unified API, usage tracking, retries/failover, custom reporting, key budgets, routing rules) with no platform fee.

    Vercelvercel.com1 minrelease
  3. Turning GLM-5.3-Flash into a Jev-like decision model

    The authors demonstrate turning an off‑the‑shelf LLM (GLM‑5.3‑Flash) into a fast, typed decision model by prompting it to output only an option index and reading the token logits. Using vLLM’s allowed_token_ids and logprob_token_ids they achieve Jev‑level accuracy and speed, even for image inputs, without any fine‑tuning.

    Hacker News front pageprivatemode.ai13 minHN13559lobste.rs2
  4. EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

    EditHero is a new benchmark for long‑horizon, part‑level 3D editing that supplies natural‑language instructions, geometry and texture targets, and a deterministic engine that produces the exact result after each edit. Using it, the authors show that bottom‑up LLM/VLM agents preserve unchanged parts better than top‑down non‑agentic methods, though they run slower (minutes per edit).

    Hugging Face Daily Papersarxiv.org1 minpaper