proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersEfstathios Karypidis, Spyros Gidaris, Nikos Komodakis1 min readpaperadvanced

Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models

Summary

Latent-Foresight is an end-to-end framework that jointly learns a latent tokenizer and a flow-based generative dynamics model for world modeling. It explicitly shapes representations for temporal predictability, outperforming two-stage baselines in future scene understanding tasks.

  • Traditional world models often use two-stage pipelines, decoupling representation learning from temporal prediction.
  • This decoupling can lead to latent spaces not optimally structured for predictable dynamics.
  • Latent-Foresight jointly learns a latent tokenizer and a flow-based generative dynamics model.
  • Key design choices prevent latent collapse and align reconstruction with generative objectives for stable optimization.

This work is important for researchers and practitioners in AI and robotics seeking more robust and efficient methods for predicting future scene evolution in complex environments.

7/10

Related reading

  1. Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

    This paper introduces Movement Trend Guidance (MTG), a method to provide foresight to 3D diffusion policies for robotic manipulation without explicit trajectory planning. MTG learns a compact latent representation of interaction evolution, significantly improving performance on various benchmarks with minimal parameter overhead.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    The paper pinpoints two weaknesses in latent‑reasoning pipelines—suboptimal visual token encoders and deterministic RL sampling—and proposes Scaffolding Minds, which trains a dedicated scaffolding encoder and a stochastic RL sampler (learning both mean and variance). This yields up to +19 points improvement on spatial planning tasks and +5.6 points average gain on nine visual reasoning benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

    Latent Interface Training (LIT) first learns a goal‑conditioned action prior without visual input, then adds a pose‑supervised latent interface as the only visual conditioning path. Applied to several vision‑language‑action models, LIT cuts vision‑action shortcuts and lifts LIBERO‑Plus success by 3.9–10.7 points and real‑world task success by 13.3–16.7 points under distribution shifts.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. LastOPD: Taming Collapse in Latent On-Policy Distillation

    The authors find that latent supervision in on‑policy distillation can cause a dramatic performance collapse even as alignment improves. Their LastOPD method limits latent alignment to the final‑layer state and crossfades to token‑level supervision, delivering 4–5 point MATH‑500 gains and halving the steps to match token‑only OPD.

    Hugging Face Daily Papersarxiv.org2 minpaper
  5. Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

    PoS is an inference-time framework that constructs and maintains explicit belief states for LLM agents, combining world state and unresolved task requirements. It validates consistency and detects "Belief Trapping" to ensure progress, achieving superior performance on long-horizon execution and diagnosis benchmarks across multiple LLM backbones.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. D-JEPA: A Decision-Aligned Latent World Model

    D-JEPA is a latent world model designed to bridge the gap between predicted outcomes and actual decision success in robotics. It learns decision-relevant relationships from executed actions, improving action selection by aligning latent space geometry with real-world results.

    Hugging Face Daily Papersarxiv.org1 minpaper