Hugging Face Daily PapersLinquan Wu, Shichang Meng, Tianxiang Jiang1 min readpaperadvanced
Draft-KV: Learning Useful Latent Communication Between Language Models
Summary
Draft‑KV introduces a lightweight interface that passes a sharer LLM’s key‑value cache to a frozen receiver via gated attention, training only 1.05 M parameters. This latent communication lifts a 0.5 B receiver from ~37 % to 78 % on MMLU‑Redux and scales with the sharer size, while the sharer can be omitted at inference.
- Latent communication can boost a frozen receiver’s performance without needing the sharer at inference; swapping messages changes accuracy by ≤0.60 points.
- Draft‑KV transmits the sharer’s KV cache through a gated‑attention side memory, training only 1.05 M parameters (≈0.3 % of C2C).
- Progressive training moves from message reconstruction to answer supervision, guarded against harmful mismatched messages.
- Scaling the sharer from 0.6 B to 8 B raises receiver accuracy on MMLU‑Redux from 46 % to 78 %.
LLM engineers and researchers building multi‑model pipelines should care because it shows a cheap way to combine models and retain gains without runtime dependence on the larger model.
8/10