proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersLinquan Wu, Shichang Meng, Tianxiang Jiang1 min readpaperadvanced

Draft-KV: Learning Useful Latent Communication Between Language Models

Summary

Draft‑KV introduces a lightweight interface that passes a sharer LLM’s key‑value cache to a frozen receiver via gated attention, training only 1.05 M parameters. This latent communication lifts a 0.5 B receiver from ~37 % to 78 % on MMLU‑Redux and scales with the sharer size, while the sharer can be omitted at inference.

  • Latent communication can boost a frozen receiver’s performance without needing the sharer at inference; swapping messages changes accuracy by ≤0.60 points.
  • Draft‑KV transmits the sharer’s KV cache through a gated‑attention side memory, training only 1.05 M parameters (≈0.3 % of C2C).
  • Progressive training moves from message reconstruction to answer supervision, guarded against harmful mismatched messages.
  • Scaling the sharer from 0.6 B to 8 B raises receiver accuracy on MMLU‑Redux from 46 % to 78 %.

LLM engineers and researchers building multi‑model pipelines should care because it shows a cheap way to combine models and retain gains without runtime dependence on the larger model.

8/10

Related reading

  1. WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

    WhiteMatter introduces all-to-all cross-layer connections in Transformers by mixing past-token representations from any depth into shared KV cache channels. This approach allows for performance comparable to 50% larger standard Transformers or improved performance with half the KV cache size. It also addresses training slowdowns with a novel cyclic iteration method.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

    Flash-dLLM is a training-free framework that accelerates Diffusion LLM inference by addressing GPU memory I/O bottlenecks with an I/O-aware KV-cache kernel. It also introduces a KV-cache-driven draft-and-verify decoding strategy, achieving significant speedups (up to 11x) over prior methods.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches

    Fathom introduces a per-query read depth mechanism for sparse decoding over offloaded KV caches, allowing each query to adaptively decide how many bits of each key channel to read. This method significantly speeds up decoding for large language models with long contexts by reducing host memory traffic, achieving 1.67x faster GPU decoding on Qwen3-8B at one million tokens.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Training Compute-Optimal Large Language Models

    The paper derives a compute‑optimal scaling law showing model size and training tokens should grow together, and validates it by training a 70B‑parameter model (Chinchilla) on 1.4 T tokens that outperforms much larger LLMs.

    Hall of Famearxiv.org66 minpaper
  5. The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

    The paper surveys recent attention variants in large language models, introducing a five‑dimensional framework (Memory Representation, Update, Access, Readout, Integration) to compare them. It shows that modern LLMs increasingly treat contextual memory as a coordinated, multi‑layer resource rather than a single attention operator.

    Hugging Face Daily Papersarxiv.org1 minpaper