Hugging Face Daily PapersSiran Peng, Tianshuo Zhang, Tianyu Fu1 min readpaperadvanced
VisionHOPE: Visual Backbones as Self-Modifying Learning Systems
Summary
VisionHOPE introduces a visual backbone that updates its own parameters on‑the‑fly using five coupled memories, with a stability‑matched step‑size scheme that keeps updates non‑expansive. The approach attains competitive ImageNet, COCO and ADE20K performance, showing self‑modifying learning systems can serve as practical general‑purpose vision models.
- VisionHOPE treats the backbone as a self-modifying learning system with five interacting memory matrices (content, key, value, learning‑rate, retention).
- Stability is enforced via a soft cap on self‑referential injection and a spectral norm clamp on memory transition, guaranteeing non‑expansive dynamics per scan.
- Images are processed by scanning rows and columns in four directions, aligning Nested Learning chunks with image rows/columns.
- Empirically matches or exceeds standard backbones on ImageNet‑1K, COCO detection, and ADE20K segmentation.
Researchers and engineers building next‑generation vision models should care because VisionHOPE shows a principled way to make backbones adapt per input while preserving stability, opening a new design space for adaptive visual systems.
8/10