1
When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
The authors show that training LLM reviewers on synthetic reviews leads to a compression of rating distributions and loss of semantic diversity, a phenomenon they call scientific-judgment collapse. They mitigate it with TrustReviewer, which uses curated training data and activation steering to preserve judgment diversity.
Hugging Face Daily Papersarxiv.org1 minpaper
