Hugging Face Daily PapersZongxia Li, Yucheng Shi, Zhongzhi Li1 min readpaperadvanced
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
Summary
This paper introduces Recursive Self-Rewrite (RSR), a framework that enables a single base LLM to discover solutions for complex tasks using diverse specialized environments (harnesses). It then reconstructs these successful trajectories into training data suitable for a general environment, significantly improving the model's performance on various benchmarks.
- RSR uses a base LLM to solve complex tasks by leveraging multiple specialized execution harnesses.
- Successful solutions found in diverse harnesses are rewritten into training trajectories for a general harness.
- The framework includes a planner to extract procedures, a critic for quality control, and an executor for execution.
- Training on these self-rewritten trajectories significantly outperforms the base model and direct trajectory SFT.
Engineers working on improving LLM capabilities for complex, multi-step tasks can use this framework to leverage specialized tools during development and distill that knowledge into a generally deployable model.
8/10