proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersHongjun Liu, Chen Zhao1 min readpaperadvanced

Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term Agents

Summary

Long-running LLM agents face a "distributed-evidence paradox" where compressed memories may not be fully supported by interaction history. The DerivAudit framework shows that while broader context can validate many memories, a significant portion remains unsupported, highlighting challenges in reliable memory admission.

  • LLM agents' compressed memories can be unsupported by their interaction history due to scattered evidence or compositional issues.
  • The "distributed-evidence paradox" describes how valid memories may appear unsupported or composed facts lack full historical basis.
  • DerivAudit is a framework to audit if a memory is truly supported by its write-time history.
  • Auditing with broader history recovers support for nearly 60% of memories initially deemed unsupported by citations alone.

Engineers building long-term LLM agents should care about this work to ensure the integrity and reliability of agent memory, preventing agents from acting on unverified information.

8/10

Related reading

  1. Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

    PoS is an inference-time framework that constructs and maintains explicit belief states for LLM agents, combining world state and unresolved task requirements. It validates consistency and detects "Belief Trapping" to ensure progress, achieving superior performance on long-horizon execution and diagnosis benchmarks across multiple LLM backbones.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Do LLMs Have the Memory of a Goldfish?

    The article explains that LLMs don’t have persistent personal memory; all “memory” is supplied by the surrounding application via the context window, summaries, or external storage. It outlines the distinction between trained weights, working‑memory (token context), and persistent application memory, shows how to construct API calls to preserve conversation state, and discusses the cost and laten…

    ByteByteGobytebytego.com12 min
  3. Shared Selective Persistent Memory for Agentic LLM Systems

    Apple proposes a memory architecture for agentic LLMs that selectively persists reusable context (specs, schemas, configs, constraints) across sessions and users. Shared workspaces with role‑based access and a zero‑token data‑refresh mechanism cut token usage by 97×, reduce task time by 14×, and raise task‑completion rates to 96% versus 71%‑79% for baselines.

    Apple Machine Learning Researchapple.com1 minpaper
  4. Verifiable Social Reasoning for LLM Assistants

    The paper introduces Fuse, a multi‑agent simulation that gives LLM assistants a verifiable ground‑truth task for social reasoning by hiding a target agent’s motive and letting a user‑mediated conversation infer it. Experiments on 12 LLMs show user mediation makes reasoning harder, models are biased by user framing, need more detail than humans, and longer chats don’t always help.

    Hugging Face Daily Papersarxiv.org1 minpaper