Related reading
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
The paper presents RecreationWorld, a five‑platform framework that lets hybrid computer‑use agents learn by recreating the behavior of a running reference, and introduces RecreationBench, a 250‑task benchmark with programmatic and visual assertions. Experiments show GPT‑6 Astra reaches 58.1% overall but struggles with deeper programmatic tests, highlighting gaps in current agents.
Hugging Face Daily Papersarxiv.org2 minpaperProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
ProgramDistill is a new benchmark that automatically extracts 1,975 replay‑verified feature behaviors from 26 real web apps, builds 4,063 coding‑agent tasks, and measures how well state‑of‑the‑art agents (e.g., GPT‑6 Astra, Claude Opus 5) can reconstruct full or partial applications, revealing steep drops in success as restoration depth grows.
Hugging Face Daily Papersarxiv.org1 minpaperShared Selective Persistent Memory for Agentic LLM Systems
Apple proposes a memory architecture for agentic LLMs that selectively persists reusable context (specs, schemas, configs, constraints) across sessions and users. Shared workspaces with role‑based access and a zero‑token data‑refresh mechanism cut token usage by 97×, reduce task time by 14×, and raise task‑completion rates to 96% versus 71%‑79% for baselines.
Apple Machine Learning Researchapple.com1 minpaper

