proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersWenbo Zhang, Xiang Ren1 min readpaperadvanced

WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

Summary

WhiteMatter introduces all-to-all cross-layer connections in Transformers by mixing past-token representations from any depth into shared KV cache channels. This approach allows for performance comparable to 50% larger standard Transformers or improved performance with half the KV cache size. It also addresses training slowdowns with a novel cyclic iteration method.

  • WhiteMatter allows Transformer layers to access past-token representations from any depth, not just their own.
  • A learned mixer combines these representations into shared key-value (KV) cache channels.
  • This can reduce KV cache size while maintaining or improving model performance.
  • With a full-size cache, WhiteMatter performs comparably to a standard Transformer with 50% more layers.

This paper offers a significant architectural modification for Transformers, providing a path to more efficient LLMs through better information reuse and KV cache optimization, which is critical for inference costs.

8/10

Related reading

  1. Draft-KV: Learning Useful Latent Communication Between Language Models

    Draft‑KV introduces a lightweight interface that passes a sharer LLM’s key‑value cache to a frozen receiver via gated attention, training only 1.05 M parameters. This latent communication lifts a 0.5 B receiver from ~37 % to 78 % on MMLU‑Redux and scales with the sharer size, while the sharer can be omitted at inference.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

    Flash-dLLM is a training-free framework that accelerates Diffusion LLM inference by addressing GPU memory I/O bottlenecks with an I/O-aware KV-cache kernel. It also introduces a KV-cache-driven draft-and-verify decoding strategy, achieving significant speedups (up to 11x) over prior methods.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

    The paper demonstrates that transformer LLMs exhibit a linear superposition property where combined inputs produce a blended next‑token distribution, and that lightweight fine‑tuning can restore this linearity. It also introduces a guided decoding algorithm that extracts two distinct, coherent continuations from one forward pass.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Transformers now runs llama.cpp quants

    Transformers now supports GGUF quantized checkpoints (used by llama.cpp) via the `kernels` library. Load a GGUF model with `from_pretrained` on Apple Silicon, run generation with standard Transformers APIs, or serve it via `transformers serve`. Benchmarks on an M2 Max show token‑throughput comparable to llama.cpp, with fallback to dequantization if kernels are missing. The integration also enable…

    Hugging Facehuggingface.co10 minHN4
  5. Scaling Laws for Looped Mixture of Experts

    This paper introduces Loop Scaling Laws, which jointly model recurrence and sparsity in Mixture-of-Experts (MoE) transformers. It finds that recurrence and sparsity offer complementary efficiency gains, improving prediction accuracy and enabling more efficient large model design.

    Hugging Face Daily Papersarxiv.org1 minpaperHN2