proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersDingyuan Dai, Heli Qi, Lei Liu2 min readpaperadvanced

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

Summary

OSWorld-Science introduces a benchmark suite of 146 scientific software tasks for evaluating visual language model agents across domains like molecular design and image analysis. The authors provide a harness for systematic comparison and show that current state‑of‑the‑art VLMs still struggle with many scientific workflows.

  • The benchmark covers 12 VLMs and 146 tasks spanning molecular drawing, retrosynthesis, pathology imaging, statistical computing, and physical simulation.
  • Task evaluation is execution‑based, inspecting application state and generated artifacts (e.g., molecular structures, segmentation masks, plots) with partial credit for incomplete results.
  • A unified harness integrates model adapters, interaction‑loop control, and trajectory logging to enable fair comparisons of agents and interaction strategies.
  • Empirical results reveal that even top VLMs with a strong harness fail on many scientific tasks, highlighting gaps in reasoning, context handling, and multi‑lingual support.

Researchers building AI agents for scientific software need a realistic, verifiable benchmark to measure progress and guide future model and harness improvements.

6/10

Related reading

  1. ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

    ScienceIDE is a framework that turns existing scientific software repositories into programmable environments that agents can use for task generation, execution, and verification. Training on these environments yields LLMs (PhAI‑IDE series) that outperform baselines on scientific code repair and several general code‑reasoning benchmarks.

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design

    CADWorld is a new benchmark suite of 200 long‑horizon mechanical CAD tasks in FreeCAD, covering sketching, part modeling, assembly, CAM, FEM, and more. Agents interact via screenshots and GUI actions; success is checked by executable validation of the saved CAD artifacts. Seven existing agents achieve at most 17.5 % success versus an 87 % expert baseline, highlighting the gap between GUI competen…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

    The paper presents RecreationWorld, a five‑platform framework that lets hybrid computer‑use agents learn by recreating the behavior of a running reference, and introduces RecreationBench, a 250‑task benchmark with programmatic and visual assertions. Experiments show GPT‑6 Astra reaches 58.1% overall but struggles with deeper programmatic tests, highlighting gaps in current agents.

    Hugging Face Daily Papersarxiv.org2 minpaper
  4. ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

    ScienceBuddy is an interactive workspace that converts researcher prompts, feedback, and execution traces into continual‑learning tasks for AI agents. It introduces a "recursive‑in‑recursive" self‑improvement loop that alternates harness refinement and model training, and showcases case studies across four scientific task families.

    Hugging Face Daily Papersarxiv.org1 minpaperHN2