proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersLizhou Liang, Xinyu Zhong, Miao Pan1 min readpaperadvanced

EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

Summary

EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.

  • EMem‑Bench contains 2,554 episodes across four task families that require agents to build, update, and reuse memory to solve later tasks.
  • Four memory failure modes are identified: fine‑grained visual recall, dynamic world‑state tracking, recording interaction outcomes, and generalization.
  • EMem external memory splits experience into spatial, event, and scene memories, exposing a unified API for policy use.
  • EMem‑8B, an 8‑billion‑parameter policy, leverages EMem and outperforms baseline models on all four challenges.

Anyone building or evaluating long‑horizon embodied agents or multimodal LLMs needs a concrete way to measure memory capabilities, and EMem‑Bench provides that yardstick.

8/10

Related reading

  1. VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

    VA‑Bench is a new benchmark that evaluates general‑purpose multimodal LLMs on the full observe‑reason‑act‑revise loop in embodied robotics, using RGB demonstrations, active camera control, and metric Cartesian commands. The best model reaches 53.9% average task success, showing active perception helps but long‑horizon tasks remain unsolved.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

    APM-Bench is a new benchmark for evaluating persistent memory in egocentric streaming video assistants across intermittent sessions. It reveals a significant utility-latency-storage trade-off, showing current models struggle with long-term recall, low overhead, and proactive assistance simultaneously.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

    HarnessVLN introduces a zero‑shot, training‑free embodied navigation framework that wraps a multimodal LLM in an "Agent Harness" – a tool‑based protocol that validates planner actions against spatial evidence, tracks progress with hierarchical event memory, and maintains a persistent spatiotemporal graph for recovery. The system works for instruction‑following and object‑goal tasks, achieving 60.…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

    PoS is an inference-time framework that constructs and maintains explicit belief states for LLM agents, combining world state and unresolved task requirements. It validates consistency and detects "Belief Trapping" to ensure progress, achieving superior performance on long-horizon execution and diagnosis benchmarks across multiple LLM backbones.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. NavHarness: Towards Lifelong Embodied Navigation

    NavHarness is a training-free system for lifelong embodied navigation that integrates memory processing into the agent's reasoning loop. It significantly improves task success rates on benchmarks by leveraging evolving maps, task records, and house knowledge across successive tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper