proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersZhen Wang, Changpeng Wang, Zhe Liu2 min readpaperadvanced

PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation

Summary

PanoVLN is a novel approach for Vision-and-Language Navigation that effectively utilizes panoramic observations. It achieves significant performance gains by modifying action prediction, training supervision, and visual representation, outperforming prior SOTA on R2R-CE and RxR-CE benchmarks.

  • Panoramic observations in VLN require specific architectural and training adjustments beyond simple image replacement for effective use.
  • PanoVLN introduces Confidence-Guided Execution (CGE) for dynamic, longer-horizon action planning from a single panoramic view.
  • Training data for panoramic VLN should include frequent branching points and clear instructions to improve route selection capabilities.
  • Combining semantic and geometric features from RGB panoramas is crucial for understanding spatial relationships in wide-field views.

This work is important for researchers and engineers developing autonomous navigation systems, as it significantly advances the state-of-the-art in vision-and-language navigation using panoramic inputs.

8/10

Related reading

  1. HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

    HarnessVLN introduces a zero‑shot, training‑free embodied navigation framework that wraps a multimodal LLM in an "Agent Harness" – a tool‑based protocol that validates planner actions against spatial evidence, tracks progress with hierarchical event memory, and maintains a persistent spatiotemporal graph for recovery. The system works for instruction‑following and object‑goal tasks, achieving 60.…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

    OmniSeek is a framework that turns a large language model into a multi‑turn audio‑visual reasoning agent by letting it decide when to look or listen and which temporal windows to fetch as evidence. The authors train it on a synthetic 170K trajectory dataset, then refine with reinforcement learning and an Audio‑Visual Necessity loss, reporting consistent gains on several benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper