Hugging Face Daily PapersLizhou Liang, Xinyu Zhong, Miao Pan1 min readpaperadvanced
EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks
Summary
EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.
- EMem‑Bench contains 2,554 episodes across four task families that require agents to build, update, and reuse memory to solve later tasks.
- Four memory failure modes are identified: fine‑grained visual recall, dynamic world‑state tracking, recording interaction outcomes, and generalization.
- EMem external memory splits experience into spatial, event, and scene memories, exposing a unified API for policy use.
- EMem‑8B, an 8‑billion‑parameter policy, leverages EMem and outperforms baseline models on all four challenges.
Anyone building or evaluating long‑horizon embodied agents or multimodal LLMs needs a concrete way to measure memory capabilities, and EMem‑Bench provides that yardstick.
8/10