proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYu Luo, Jiamin Jiang, Yimin Zuo1 min readpaperadvanced

Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

Summary

PoS is an inference-time framework that constructs and maintains explicit belief states for LLM agents, combining world state and unresolved task requirements. It validates consistency and detects "Belief Trapping" to ensure progress, achieving superior performance on long-horizon execution and diagnosis benchmarks across multiple LLM backbones.

  • PoS creates and maintains explicit belief states for LLM agents to ensure a coherent understanding of the world.
  • Belief states integrate current world state estimates with unresolved task requirements, making explicit what the agent needs to learn.
  • The framework validates belief consistency and monitors task progress to detect "Belief Trapping," where agents act without progress.
  • Tailored recovery mechanisms are applied based on trapping patterns and unresolved task requirements.

This framework provides a robust method for LLM agents to manage complex, long-horizon tasks, which is critical for developing more reliable and effective AI agents.

8/10

Related reading

  1. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Shared Selective Persistent Memory for Agentic LLM Systems

    Apple proposes a memory architecture for agentic LLMs that selectively persists reusable context (specs, schemas, configs, constraints) across sessions and users. Shared workspaces with role‑based access and a zero‑token data‑refresh mechanism cut token usage by 97×, reduce task time by 14×, and raise task‑completion rates to 96% versus 71%‑79% for baselines.

    Apple Machine Learning Researchapple.com1 minpaper
  3. Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems

    Emergence World is a continuously running multi‑agent sandbox used to stress‑test frontier LLM‑based agents over weeks. Eight parallel worlds (seven homogeneous, one mixed) generated 850 k LLM calls and ~50 B tokens while agents pursued goals, used tools, and maintained persistent memory. The authors injected three adversarial events—prompt injection, misinformation, and private‑memory exposure—a…

    Hugging Face Daily Papersarxiv.org1 minpaper