Hugging Face Daily PapersYiran Wang, Xingyilang Yin, Junfu Pu1 min readpaperadvanced
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Summary
The paper presents GameHorizon, a unified suite comprising an automated annotation pipeline, a 5,000‑hour multi‑horizon gameplay dataset from 21 AAA titles, and reproducible offline and online benchmarks. Using it, the authors evaluate 47 models, exposing a hierarchy of task difficulty and gaps in long‑term planning.
- GameHorizon‑Annotator automates creation of multi‑horizon language instructions aligned with video and action streams.
- GameHorizon‑Data offers 5,000 h of AAA gameplay from 21 titles with temporally aligned video, actions, and instructions.
- GameHorizon‑Bench supplies reproducible offline QA and stepwise online evaluation that links scores to real gameplay performance.
- Benchmarking 47 models reveals a clear hierarchy of task difficulty and gaps in long‑horizon planning.
Game AI and multimodal model researchers should care because it provides the first large‑scale, multi‑horizon benchmark for reliably comparing gameplay capabilities across models.
7/10
