Hugging Face Daily PapersSeongtae Hong, Youngjoon Jang, Jungseob Lee1 min readpaperadvanced
RenderRank: Learning to Rerank Text with Compressed Visual Tokens
Summary
RenderRank renders documents as images and uses a vision‑language model to produce compressed visual tokens for reranking, cutting input length by up to 35% while achieving higher NDCG@10 than text‑only baselines. It shows especially strong gains on long‑document datasets with half the token count and 1.7× throughput.
- Rendering text to images reduces input tokens by 16.5‑35.5% versus raw text for each candidate document.
- On BEIR (11 datasets), RenderRank attains avg NDCG@10 = 55.96, beating all text‑only rerankers under 4B parameters.
- For long‑document sets, it reaches avg NDCG@10 = 88.27 with ~50% fewer tokens and 1.70× higher throughput.
- Training uses two‑stage distillation: align visual scores to a text teacher, then fine‑tune with query‑specific positives/negatives.
IR engineers building rerankers for large corpora or long documents should care because visual token encoding cuts latency and memory while preserving or improving relevance quality.
7/10