proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersSeongtae Hong, Youngjoon Jang, Jungseob Lee1 min readpaperadvanced

RenderRank: Learning to Rerank Text with Compressed Visual Tokens

Summary

RenderRank renders documents as images and uses a vision‑language model to produce compressed visual tokens for reranking, cutting input length by up to 35% while achieving higher NDCG@10 than text‑only baselines. It shows especially strong gains on long‑document datasets with half the token count and 1.7× throughput.

  • Rendering text to images reduces input tokens by 16.5‑35.5% versus raw text for each candidate document.
  • On BEIR (11 datasets), RenderRank attains avg NDCG@10 = 55.96, beating all text‑only rerankers under 4B parameters.
  • For long‑document sets, it reaches avg NDCG@10 = 88.27 with ~50% fewer tokens and 1.70× higher throughput.
  • Training uses two‑stage distillation: align visual scores to a text teacher, then fine‑tune with query‑specific positives/negatives.

IR engineers building rerankers for large corpora or long documents should care because visual token encoding cuts latency and memory while preserving or improving relevance quality.

7/10

Related reading

  1. Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion

    D‑RAC extends the W‑RAC pipeline to any document type by first rendering it to PDF, then using a single multimodal LLM pass to turn each page into a retrieval‑optimized Markdown representation (tables become prose, headings are kept). Chunking works on deterministic IDs, avoiding re‑tokenising the source text. On a 236‑doc, 795‑page benchmark D‑RAC processes the whole set in 72 min, creates 1,748…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

    This paper introduces a training-adaptive Convolutional Sparse Coding (CSC) framework where the sparsity coefficient is learned end-to-end via FISTA unfolding. It uses an information bottleneck perspective to balance representation compression and content preservation, showing improved robustness to input perturbations on CIFAR and ImageNet.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

    FLAT is a pre‑training framework that encodes images and text into a shared 1D token sequence with variable length via nested dropout, enabling the same embeddings for cross‑modal retrieval and generation. It attains state‑of‑the‑art scores on COCO and Flickr30K for captioning, retrieval, and text‑to‑image generation, and supports interpolation and zero‑shot composed retrieval.

    Hugging Face Daily Papersarxiv.org1 minpaper