Related reading
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
The paper presents GameHorizon, a unified suite comprising an automated annotation pipeline, a 5,000‑hour multi‑horizon gameplay dataset from 21 AAA titles, and reproducible offline and online benchmarks. Using it, the authors evaluate 47 models, exposing a hierarchy of task difficulty and gaps in long‑term planning.
Hugging Face Daily Papersarxiv.org1 minpaperAI for Games in the Foundation Model Era
The paper surveys how foundation models are used across six roles in the game development lifecycle—from playing agents to design assistance and runtime adaptation. It highlights limited transferability due to game-specific interfaces and notes that evaluation is mature for bounded play but weak for adaptive and testing scenarios.
Hugging Face Daily Papersarxiv.org1 minpaperHow good are frontier models at physics?
The authors audit six popular physics benchmarks by having domain experts re‑grade model outputs, fixing reference answers and removing ambiguous items. After correction, GPT‑5.6‑Sol’s mean@4 jumps from ~47 % to ~79 % on HLE‑Physics and from ~61 % to ~87 % on CMT‑Benchmark, with a corrected pass@4 of 94 % on 54 vetted CritPt challenges. The work shows current benchmarks severely under‑report LLM…
Hacker News front pagearxiv.org2 minpaperHN9650Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks
Android Bench 2.0 adds a set of long‑horizon tasks (multi‑day Android development problems) and introduces agent‑based evaluation. Scoring is now continuous, with the best model achieving a 28 % pass rate on these tasks, far lower than the ~91 % on earlier short tasks. The post lists new models on the leaderboard and points to updated methodology and GitHub repo.
Article: Beyond Relevance: A Governance-First Architecture for Enterprise Personalization
The article proposes a governance‑first architecture for enterprise personalization, where policy‑driven steps (memory, journey graph, AI routing, scoring, trust checks, outcome simulation) shape the recommendation before it is returned. A reference FastAPI implementation demonstrates the pattern with external YAML policies and optional LLM assistance.
InfoQinfoq.com19 min


