1
Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches
Fathom introduces a per-query read depth mechanism for sparse decoding over offloaded KV caches, allowing each query to adaptively decide how many bits of each key channel to read. This method significantly speeds up decoding for large language models with long contexts by reducing host memory traffic, achieving 1.67x faster GPU decoding on Qwen3-8B at one million tokens.
Hugging Face Daily Papersarxiv.org1 minpaper
