proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersXin Li, Mengbing Liu, Chau Yuen1 min readpaperadvanced

Measuring Collapse and Correction in Homogeneous-Panel LLM Debate

Summary

The authors present an auditable protocol that records the dynamics of homogeneous multi‑agent LLM debates on multiple‑choice questions, tracking collapse, correction, and intervention utility. On 6,925 MMLU‑Pro debates they show that preventing collapses can also suppress useful corrections, and that early‑round disagreement predicts collapse risk.

  • The protocol logs each debate as a transition ledger of collapse, correction, onset, and signed intervention utility, enabling fine‑grained audit beyond final accuracy.
  • In 6,925 MMLU‑Pro debates 253 collapses were identified; a leave‑one‑model‑out probe‑gated freeze stops 29 collapses but eliminates 108 corrections, exposing a trade‑off.
  • An 8‑probe pre‑debate screen correlates strongly (Spearman ρ=0.893) with conditional‑collapse risk, though it is not a calibrated predictor of capability.
  • Most collapses occur in the first debate round, suggesting early disagreement is a key signal for intervention design.

Researchers building multi‑agent LLM debate systems need evaluation metrics that capture both harmful collapses and beneficial corrections; this work provides a concrete framework and tooling.

8/10

Related reading

  1. Verifiable Social Reasoning for LLM Assistants

    The paper introduces Fuse, a multi‑agent simulation that gives LLM assistants a verifiable ground‑truth task for social reasoning by hiding a target agent’s motive and letting a user‑mediated conversation infer it. Experiments on 12 LLMs show user mediation makes reasoning harder, models are biased by user framing, need more detail than humans, and longer chats don’t always help.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

    The paper introduces BI‑Bench, a new benchmark of real‑world BI questions derived from public dashboards, and BI‑Agent, a tool‑augmented LLM system that breaks BI workflows into search, join, and transform subtasks. Baseline LLMs hit <50 % accuracy on BI‑Bench. By orchestrating specialized data‑management tools and post‑training the model with supervised fine‑tuning and reinforcement learning on…

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. Constraint Decay: The Fragility of LLM Agents in Backend Code Generation

    A systematic evaluation of LLM agents generating multi‑file backend code shows a sharp drop in correctness when structural constraints (framework conventions, ORM usage, API contracts) are added. Across 100 tasks in 8 Python web frameworks, assertion pass rates fall ~27 points, with data‑layer bugs (bad queries, ORM violations) driving most failures. Mid‑size models cope with minimal frameworks (…

    arXiv cs.SE (Software Engineering)arxiv.org1 minpaperHN287197
  4. When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

    The authors show that training LLM reviewers on synthetic reviews leads to a compression of rating distributions and loss of semantic diversity, a phenomenon they call scientific-judgment collapse. They mitigate it with TrustReviewer, which uses curated training data and activation steering to preserve judgment diversity.

    Hugging Face Daily Papersarxiv.org1 minpaper