Hugging Face Daily PapersYanlong Chen, Yining Chen, Song Zhang1 min readpaperadvanced
PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers
Summary
PrismQuant proposes a quantizer-aware rotation that aligns activation eigenspaces with grouped INT4 subspaces, yielding a closed‑form optimal solution. Experiments on Llama, Qwen, and Mistral show state‑of‑the‑art perplexity/accuracy and notable speed‑memory gains in real deployments.
- Formulates rotation design as a Ky Fan trace maximization and provides a provably optimal closed‑form solution.
- Uses compact Householder transforms for gradient‑free construction, applicable both offline (foldable) and online.
- Achieves SOTA W4A4KV4 results: Llama‑3.2‑3B best perplexity/accuracy, Llama‑3.1‑70B within 0.22 % of FP precision.
- Deployment on Llama‑3.1‑8B gives 1.51× prefill and 1.22× decode speedups vs FP16, with 56 % lower peak memory.
LLM engineers looking to cut inference cost with low‑bit quantization will benefit from a rigorously optimal rotation that improves accuracy and runtime without extra training.
8/10