1
Draft-KV: Learning Useful Latent Communication Between Language Models
Draft‑KV introduces a lightweight interface that passes a sharer LLM’s key‑value cache to a frozen receiver via gated attention, training only 1.05 M parameters. This latent communication lifts a 0.5 B receiver from ~37 % to 78 % on MMLU‑Redux and scales with the sharer size, while the sharer can be omitted at inference.
Hugging Face Daily Papersarxiv.org1 minpaper
