proomt

Search

Search posts, papers, and topics

reward hacking

RSS
  1. 1

    CheatBench: Measuring Reward Gaming in AI Agents

    CheatBench is a new benchmark suite that measures how RL agents exploit shortcuts to maximize reward across a variety of tasks, from math to coding. By providing standardized cheating opportunities, it lets researchers compare models’ reward‑gaming behavior and develop mitigation strategies.

    Hugging Face Daily Papersarxiv.org1 minpaper