Hugging Face Daily PapersAmir Taubenfeld, Zorik Gekhman, Avigail Grinstein-Dabush1 min readpaperadvanced
Verifiable Social Reasoning for LLM Assistants
Summary
The paper introduces Fuse, a multi‑agent simulation that gives LLM assistants a verifiable ground‑truth task for social reasoning by hiding a target agent’s motive and letting a user‑mediated conversation infer it. Experiments on 12 LLMs show user mediation makes reasoning harder, models are biased by user framing, need more detail than humans, and longer chats don’t always help.
- Fuse creates a controlled social‑reasoning benchmark where the target’s hidden motive provides ground truth by construction.
- A human study with 24 k annotations confirms the simulation’s fidelity to real‑world social inference.
- Evaluating 12 LLMs reveals (i) user mediation amplifies difficulty, (ii) models are sensitive to biased user framing, (iii) they often need more detail than humans to succeed, and (iv) longer dialogues don’t guarantee b…
- The authors release the Fuse framework and a 21 k‑example dataset for reproducible research.
Anyone building or evaluating LLM assistants for advice or social interaction needs a reliable, verifiable testbed for social reasoning.
7/10



