proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersXunyi Zhao, Jian Zhou, Sihao Lin1 min readpaperadvanced

NavHarness: Towards Lifelong Embodied Navigation

Summary

NavHarness is a training-free system for lifelong embodied navigation that integrates memory processing into the agent's reasoning loop. It significantly improves task success rates on benchmarks by leveraging evolving maps, task records, and house knowledge across successive tasks.

  • NavHarness integrates memory processing (maps, task records, house knowledge) into the navigation loop for lifelong learning.
  • It achieves state-of-the-art success rates on GOAT-Bench (83.7 s-SR) and IR2R-CE (85.9 s-SR) using GPT-6 Astra.
  • Structured recovery handovers are more effective than simple summaries for transferring experience between agentic sessions.
  • Consolidating experience across deployments improves navigation beyond just retaining maps and task records.

This paper is crucial for researchers and engineers developing embodied AI agents, as it addresses the fundamental challenge of continuous learning and memory management in dynamic environments.

8/10

Related reading

  1. HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

    HarnessVLN introduces a zero‑shot, training‑free embodied navigation framework that wraps a multimodal LLM in an "Agent Harness" – a tool‑based protocol that validates planner actions against spatial evidence, tracks progress with hierarchical event memory, and maintains a persistent spatiotemporal graph for recovery. The system works for instruction‑following and object‑goal tasks, achieving 60.…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

    OmniHarness introduces a symbolic‑policy framework that extracts reusable procedural knowledge from multimodal LLM‑driven visual generation runs. By decoupling task logic from instance inputs, the system can instantiate, adapt, and compose policies for new visual tasks, using intermediate verification for on‑the‑fly refinement while keeping the underlying model frozen. Self‑directed practice task…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

    The paper presents RecreationWorld, a five‑platform framework that lets hybrid computer‑use agents learn by recreating the behavior of a running reference, and introduces RecreationBench, a 250‑task benchmark with programmatic and visual assertions. Experiments show GPT‑6 Astra reaches 58.1% overall but struggles with deeper programmatic tests, highlighting gaps in current agents.

    Hugging Face Daily Papersarxiv.org2 minpaper
  5. Native Action-Prior Learning from Videos for World Action Models

    Native Action-Prior Learning (NAVA‑WAM) trains robot action policies directly from observation‑only videos by matching future video flow through a joint attention mechanism, then fine‑tunes with a small set of labeled demos. Experiments show it beats prior methods on both in‑distribution and out‑of‑distribution tasks and transfers to real robots with fewer action labels.

    Hugging Face Daily Papersarxiv.org1 minpaper