Hugging Face Daily PapersCaiqi Zhang, Rujun Han, Zifeng Wang1 min readpaperadvanced
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
Summary
VeriHarness turns an LLM generator into an agentic verifier by adding a workspace, evidence tools, and verification skills that resolve disagreement and challenge consensus. Across five long‑horizon benchmarks it improves selection scores by ~6 points over single rollouts and releases 26k rollouts for future work.
- Disagreement among multiple rollouts often reveals correct alternatives, while consensus can hide errors.
- VeriHarness equips the LLM with a workspace, evidence tools, a disagreement resolver, and a consensus challenger to verify claims.
- On five benchmarks, it outperforms baselines, adding 6.2 points (Gemini 3.5 Flash) and 6.4 points (Claude Opus 4.8) over a single rollout.
- Verification skills can self‑improve using failure feedback, enabling iterative refinement.
Anyone building LLM agents for complex, long‑horizon tasks needs a scalable way to trust outputs; VeriHarness offers a concrete, measurable solution.
8/10

