Hugging Face Daily PapersMika Okamoto, Ansel Kaplan Erol2 min readpaperintermediate
PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
Summary
PACT is a new benchmark designed to measure how well enterprise LLM agents follow compliance rules, especially when under user pressure. It found substantial variability across 22 models, with even the strongest assistants misapplying rules 6-10% of the time, and user pressure increasing violation rates by 65% on average.
- PACT evaluates LLM rule-following in 12 regulated enterprise domains and 48 multi-turn scenarios.
- The benchmark tests compliance under various 'pressures' from users or managers.
- Even top LLMs misapply rules 6-10% of the time in sensitive contexts.
- User pressure significantly increases rule violation rates by 65% on average.
Engineers deploying LLM agents in regulated enterprise environments should care, as PACT highlights critical compliance risks and the need for robust guardrails and careful model selection.
8/10



