proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersCaiqi Zhang, Rujun Han, Zifeng Wang1 min readpaperadvanced

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

Summary

VeriHarness turns an LLM generator into an agentic verifier by adding a workspace, evidence tools, and verification skills that resolve disagreement and challenge consensus. Across five long‑horizon benchmarks it improves selection scores by ~6 points over single rollouts and releases 26k rollouts for future work.

  • Disagreement among multiple rollouts often reveals correct alternatives, while consensus can hide errors.
  • VeriHarness equips the LLM with a workspace, evidence tools, a disagreement resolver, and a consensus challenger to verify claims.
  • On five benchmarks, it outperforms baselines, adding 6.2 points (Gemini 3.5 Flash) and 6.4 points (Claude Opus 4.8) over a single rollout.
  • Verification skills can self‑improve using failure feedback, enabling iterative refinement.

Anyone building LLM agents for complex, long‑horizon tasks needs a scalable way to trust outputs; VeriHarness offers a concrete, measurable solution.

8/10

Related reading

  1. Constraint Decay: The Fragility of LLM Agents in Backend Code Generation

    A systematic evaluation of LLM agents generating multi‑file backend code shows a sharp drop in correctness when structural constraints (framework conventions, ORM usage, API contracts) are added. Across 100 tasks in 8 Python web frameworks, assertion pass rates fall ~27 points, with data‑layer bugs (bad queries, ORM violations) driving most failures. Mid‑size models cope with minimal frameworks (…

    arXiv cs.SE (Software Engineering)arxiv.org1 minpaperHN287197
  2. NavHarness: Towards Lifelong Embodied Navigation

    NavHarness is a training-free system for lifelong embodied navigation that integrates memory processing into the agent's reasoning loop. It significantly improves task success rates on benchmarks by leveraging evolving maps, task records, and house knowledge across successive tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Your Agent Aced the Task. Will It Do It Again?

    The post introduces the Consistency Analyzer, a cheap black‑box diagnostic that flags flip‑prone decision steps in LLM agent traces, and shows how feeding the resulting consistency guidelines back into ALTK‑Evolve halves the gap between mean success and all‑run success (Pass⁵) on the AppWorld benchmark without hurting average accuracy.

    Hugging Facehuggingface.co8 minHN21
  4. HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses

    The paper presents HypoEvolve, a generational genetic algorithm that coordinates specialized LLM agents to iteratively propose, critique, and refine scientific hypotheses. On a drug‑repurposing benchmark across 34 cancer types, it outperforms six baselines, achieving a DepMap selectivity of 0.171 versus 0.115.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Shared Selective Persistent Memory for Agentic LLM Systems

    Apple proposes a memory architecture for agentic LLMs that selectively persists reusable context (specs, schemas, configs, constraints) across sessions and users. Shared workspaces with role‑based access and a zero‑token data‑refresh mechanism cut token usage by 97×, reduce task time by 14×, and raise task‑completion rates to 96% versus 71%‑79% for baselines.

    Apple Machine Learning Researchapple.com1 minpaper
  6. Verifiable Social Reasoning for LLM Assistants

    The paper introduces Fuse, a multi‑agent simulation that gives LLM assistants a verifiable ground‑truth task for social reasoning by hiding a target agent’s motive and letting a user‑mediated conversation infer it. Experiments on 12 LLMs show user mediation makes reasoning harder, models are biased by user framing, need more detail than humans, and longer chats don’t always help.

    Hugging Face Daily Papersarxiv.org1 minpaper