proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersZongxia Li, Yucheng Shi, Zhongzhi Li1 min readpaperadvanced

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Summary

This paper introduces Recursive Self-Rewrite (RSR), a framework that enables a single base LLM to discover solutions for complex tasks using diverse specialized environments (harnesses). It then reconstructs these successful trajectories into training data suitable for a general environment, significantly improving the model's performance on various benchmarks.

  • RSR uses a base LLM to solve complex tasks by leveraging multiple specialized execution harnesses.
  • Successful solutions found in diverse harnesses are rewritten into training trajectories for a general harness.
  • The framework includes a planner to extract procedures, a critic for quality control, and an executor for execution.
  • Training on these self-rewritten trajectories significantly outperforms the base model and direct trajectory SFT.

Engineers working on improving LLM capabilities for complex, multi-step tasks can use this framework to leverage specialized tools during development and distill that knowledge into a generally deployable model.

8/10

Related reading

  1. Recursive self-improvement of AI research agents

    The paper introduces AIDE², an AI research agent that rewrites its own code, benchmarks each version, and adopts the best performing changes—a process they call recursive self‑improvement. In an 8‑day autonomous run it produced seven improvements that beat a strong human‑engineered baseline on four unseen benchmarks and reduced reward‑hacking from 55 % to 32 %.

    Hugging Face Daily Papersarxiv.org2 minpaperHN31
  2. ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

    ROSS is a method for large language model post-training that selectively supervises historical self-generated rollouts, applying loss only to useful continuations while preserving full trajectory context. It consistently improves LLM performance across various tasks like code generation and instruction following, achieving gains on Qwen3.6-35B-A3B without requiring new policy rollouts.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. The Router Within: Eliciting Native Skill Routing from a Frozen LLM

    The paper introduces Gavel, a method that extracts a frozen LLM's internal routing signal via two trained linear maps, eliminating the need to embed skill descriptions in the prompt. Experiments on Qwen3‑32B show up to 13.4‑point improvements on task benchmarks and higher skill‑use accuracy compared to larger retrieval‑based systems.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents

    Beyond Top‑k Skill Retrieval: Diversity‑Aware Skill Routing (DSR) applies a Determinantal Point Process with a query‑residual diversity kernel to rerank skill candidates, balancing relevance and redundancy. On the SkillRouter benchmark it raises recall and full‑coverage, especially for multi‑skill queries, showing that skill routing benefits from set‑selection rather than independent ranking.

    Hugging Face Daily Papersarxiv.org1 minpaper