Hugging Face Daily PapersDingyuan Dai, Heli Qi, Lei Liu2 min readpaperadvanced
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
Summary
OSWorld-Science introduces a benchmark suite of 146 scientific software tasks for evaluating visual language model agents across domains like molecular design and image analysis. The authors provide a harness for systematic comparison and show that current state‑of‑the‑art VLMs still struggle with many scientific workflows.
- The benchmark covers 12 VLMs and 146 tasks spanning molecular drawing, retrosynthesis, pathology imaging, statistical computing, and physical simulation.
- Task evaluation is execution‑based, inspecting application state and generated artifacts (e.g., molecular structures, segmentation masks, plots) with partial credit for incomplete results.
- A unified harness integrates model adapters, interaction‑loop control, and trajectory logging to enable fair comparisons of agents and interaction strategies.
- Empirical results reveal that even top VLMs with a strong harness fail on many scientific tasks, highlighting gaps in reasoning, context handling, and multi‑lingual support.
Researchers building AI agents for scientific software need a realistic, verifiable benchmark to measure progress and guide future model and harness improvements.
6/10