Hugging Face Daily PapersWenbo Zhang, Xiang Ren1 min readpaperadvanced
WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing
Summary
WhiteMatter introduces all-to-all cross-layer connections in Transformers by mixing past-token representations from any depth into shared KV cache channels. This approach allows for performance comparable to 50% larger standard Transformers or improved performance with half the KV cache size. It also addresses training slowdowns with a novel cyclic iteration method.
- WhiteMatter allows Transformer layers to access past-token representations from any depth, not just their own.
- A learned mixer combines these representations into shared key-value (KV) cache channels.
- This can reduce KV cache size while maintaining or improving model performance.
- With a full-size cache, WhiteMatter performs comparably to a standard Transformer with 50% more layers.
This paper offers a significant architectural modification for Transformers, providing a path to more efficient LLMs through better information reuse and KV cache optimization, which is critical for inference costs.
8/10