proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersAlham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri1 min readpaperintermediate

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

Summary

HyperBrowseComp is a new multilingual and multimodal benchmark for web-browsing AI agents, featuring 423 challenging, human-validated questions across 13 languages. It requires agents to find obscure evidence and connect information from diverse sources like videos and maps, going beyond parametric knowledge.

  • HyperBrowseComp is a benchmark for evaluating web-browsing AI agents.
  • It contains 423 manually authored, human-validated questions in 13 languages.
  • Questions are designed to be extremely challenging, requiring multi-step reasoning and multimodal evidence (videos, images, maps).
  • Easier questions are filtered out by testing with models without internet access to prevent reliance on parametric knowledge.

This benchmark is crucial for developers building advanced AI agents, as it provides a rigorous, diverse testbed for evaluating their ability to perform complex, real-world information seeking across languages and media.

7/10

Related reading

  1. ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

    ProgramDistill is a new benchmark that automatically extracts 1,975 replay‑verified feature behaviors from 26 real web apps, builds 4,063 coding‑agent tasks, and measures how well state‑of‑the‑art agents (e.g., GPT‑6 Astra, Claude Opus 5) can reconstruct full or partial applications, revealing steep drops in success as restoration depth grows.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

    The paper presents RecreationWorld, a five‑platform framework that lets hybrid computer‑use agents learn by recreating the behavior of a running reference, and introduces RecreationBench, a 250‑task benchmark with programmatic and visual assertions. Experiments show GPT‑6 Astra reaches 58.1% overall but struggles with deeper programmatic tests, highlighting gaps in current agents.

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

    CodeMidas builds RL environments directly from open‑source code: agents explore a repo, infer a spec, generate tests from the original implementation, and filter tasks via execution checks. The pipeline yields 5,545 high‑quality coding tasks across 23 languages and 15 domains. Training the MiMo‑V2.5 agent with GRPO on this dataset improves benchmark scores by 8‑18% (e.g., DeepSWE +11.7%, ProgramB…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems

    Emergence World is a continuously running multi‑agent sandbox used to stress‑test frontier LLM‑based agents over weeks. Eight parallel worlds (seven homogeneous, one mixed) generated 850 k LLM calls and ~50 B tokens while agents pursued goals, used tools, and maintained persistent memory. The authors injected three adversarial events—prompt injection, misinformation, and private‑memory exposure—a…

    Hugging Face Daily Papersarxiv.org1 minpaper