proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersSiran Peng, Tianshuo Zhang, Tianyu Fu1 min readpaperadvanced

VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

Summary

VisionHOPE introduces a visual backbone that updates its own parameters on‑the‑fly using five coupled memories, with a stability‑matched step‑size scheme that keeps updates non‑expansive. The approach attains competitive ImageNet, COCO and ADE20K performance, showing self‑modifying learning systems can serve as practical general‑purpose vision models.

  • VisionHOPE treats the backbone as a self-modifying learning system with five interacting memory matrices (content, key, value, learning‑rate, retention).
  • Stability is enforced via a soft cap on self‑referential injection and a spectral norm clamp on memory transition, guaranteeing non‑expansive dynamics per scan.
  • Images are processed by scanning rows and columns in four directions, aligning Nested Learning chunks with image rows/columns.
  • Empirically matches or exceeds standard backbones on ImageNet‑1K, COCO detection, and ADE20K segmentation.

Researchers and engineers building next‑generation vision models should care because VisionHOPE shows a principled way to make backbones adapt per input while preserving stability, opening a new design space for adaptive visual systems.

8/10

Related reading

  1. Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

    This paper evaluates GPT-6 Astra and five other frontier general-purpose AI systems across 34 computer vision capabilities and 55 benchmarks. It finds these systems excel at semantic interpretation and reasoning, but struggle with metric geometric accuracy, faithful reconstruction, and fine-grained specialized knowledge.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

    Reasoning vision-language models (VLMs) are expensive due to long visual token sequences. This paper identifies a "representation-utilization gap" where pruned visual tokens are underutilized, not just lost. SCOPD, a self-distillation framework, improves VLM performance with sparse visual contexts, achieving 92.43% of unpruned performance at 10% token retention.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

    Distill frozen world‑model features into a Vision‑Language‑Action policy via a single alignment loss; no teacher at train time, no extra runtime cost. 0.8 B student runs 32 ms / 1.86 GB on RTX 5090, hits 97.9 % on LIBERO and improves RoboCasa‑GR1 from 48.2 % to 50.5 %, with transfer to real single‑arm and bimanual robots.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

    Latent Interface Training (LIT) first learns a goal‑conditioned action prior without visual input, then adds a pose‑supervised latent interface as the only visual conditioning path. Applied to several vision‑language‑action models, LIT cuts vision‑action shortcuts and lifts LIBERO‑Plus success by 3.9–10.7 points and real‑world task success by 13.3–16.7 points under distribution shifts.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

    OmniHarness introduces a symbolic‑policy framework that extracts reusable procedural knowledge from multimodal LLM‑driven visual generation runs. By decoupling task logic from instance inputs, the system can instantiate, adapt, and compose policies for new visual tasks, using intermediate verification for on‑the‑fly refinement while keeping the underlying model frozen. Self‑directed practice task…

    Hugging Face Daily Papersarxiv.org1 minpaper