proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersRui Zhong, Yu Li, Zheyu Yan1 min readpaperadvanced

Beyond Selection: Token Parameterization for Extreme Visual Token Compression

Summary

This paper introduces Braco, a lightweight coder for extreme visual token compression in vision-language models. It uses a novel token parameterization approach to maintain visual grounding and achieve high accuracy at 23x-64x compression, significantly improving efficiency over prior methods.

  • Extreme visual token compression is challenging; pruning breaks visual grounding, while learned resamplers add complexity and cost.
  • Braco proposes token parameterization, separating basis transformation/truncation from coordinate organization for compression.
  • The four-step coder combines transform-basis truncation, input-independent basis-coordinate embeddings, and budget-dependent orthogonal re-parameterization.
  • Braco achieves 95.2% accuracy at 23x-64x compression, remaining competitive at 144x, forming a favorable accuracy-efficiency frontier.

Engineers working with vision-language models will find this relevant for significantly improving model efficiency and reducing computational costs under extreme compression budgets.

8/10

Related reading

  1. RenderRank: Learning to Rerank Text with Compressed Visual Tokens

    RenderRank renders documents as images and uses a vision‑language model to produce compressed visual tokens for reranking, cutting input length by up to 35% while achieving higher NDCG@10 than text‑only baselines. It shows especially strong gains on long‑document datasets with half the token count and 1.7× throughput.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

    Reasoning vision-language models (VLMs) are expensive due to long visual token sequences. This paper identifies a "representation-utilization gap" where pruned visual tokens are underutilized, not just lost. SCOPD, a self-distillation framework, improves VLM performance with sparse visual contexts, achieving 92.43% of unpruned performance at 10% token retention.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

    The paper presents SemanTok, a flexible video tokenizer that injects frozen DINO features and reconstructs them from any token prefix, achieving strong semantic alignment and fidelity. A 201 M SemanTok AR model matches or exceeds a 3.4× larger VideoFlexTok baseline, with cheaper short‑prefix prediction and better generation quality.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. tokenizers v1: encode, decode and scaling, measured

    Hugging Face has released `tokenizers` v1, a major performance update that achieves 3-30x faster encoding than v0.23 while maintaining identical output and API compatibility. Key optimizations include a SIMD-accelerated splitter, a thread-local word cache, and an allocation-free BPE merge loop, ensuring tokenization doesn't bottleneck ML workflows.

    Hugging Facehuggingface.co10 min
  5. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

    DeepSeek‑V4.1‑Flash is a 552B‑parameter multimodal Mixture‑of‑Experts LLM that supports up to 1 M‑token contexts while slashing KV‑cache memory to 890 bytes/token (≈¼ of its predecessor) via cross‑layer reuse (CSA2) and FP4 quantisation, plus a SWA‑Bounded Replay scheme that cuts persistent cache to 1/8. The Causal Encoder‑Decoder design halves prefill compute (8B vs 16B active parameters) and th…

    Hugging Face Daily Papersarxiv.org3 minpaperHN12710
  6. Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

    This paper introduces a training-adaptive Convolutional Sparse Coding (CSC) framework where the sparsity coefficient is learned end-to-end via FISTA unfolding. It uses an information bottleneck perspective to balance representation compression and content preservation, showing improved robustness to input perturbations on CIFAR and ImageNet.

    Hugging Face Daily Papersarxiv.org1 minpaper