Hugging Face Daily PapersAhmadreza Jeddi, Enming Zhang, Jasper Gerigk1 min readpaperadvanced
SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models
Summary
Reasoning vision-language models (VLMs) are expensive due to long visual token sequences. This paper identifies a "representation-utilization gap" where pruned visual tokens are underutilized, not just lost. SCOPD, a self-distillation framework, improves VLM performance with sparse visual contexts, achieving 92.43% of unpruned performance at 10% token retention.
- VLM inference is costly due to processing images/videos as long sequences of visual tokens.
- Performance degradation from token pruning is partly due to underutilization of remaining information, termed the "representation-utilization gap".
- SCOPD uses on-policy self-distillation with a full-context teacher to train a student on pruned visual tokens.
- SCOPD+ enhances this by selectively distilling visually sensitive response positions with a small visual-budget intervention.
Engineers and researchers deploying vision-language models will find this method useful for significantly reducing inference costs while maintaining high performance.
8/10