proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYanlong Chen, Yining Chen, Song Zhang1 min readpaperadvanced

PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers

Summary

PrismQuant proposes a quantizer-aware rotation that aligns activation eigenspaces with grouped INT4 subspaces, yielding a closed‑form optimal solution. Experiments on Llama, Qwen, and Mistral show state‑of‑the‑art perplexity/accuracy and notable speed‑memory gains in real deployments.

  • Formulates rotation design as a Ky Fan trace maximization and provides a provably optimal closed‑form solution.
  • Uses compact Householder transforms for gradient‑free construction, applicable both offline (foldable) and online.
  • Achieves SOTA W4A4KV4 results: Llama‑3.2‑3B best perplexity/accuracy, Llama‑3.1‑70B within 0.22 % of FP precision.
  • Deployment on Llama‑3.1‑8B gives 1.51× prefill and 1.22× decode speedups vs FP16, with 56 % lower peak memory.

LLM engineers looking to cut inference cost with low‑bit quantization will benefit from a rigorously optimal rotation that improves accuracy and runtime without extra training.

8/10

Related reading

  1. Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

    Pruned CTC limits CTC alignment calculations to the batch‑specific subset of tokens actually needed, eliminating the linear memory blow‑up with large vocabularies while preserving exact loss and gradient values. Applied to LLM‑based ASR, it cuts per‑step memory 5.1× with modest compute overhead and delivers near‑baseline WER at 7‑10× faster inference for both offline and bounded‑history streaming.

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. tokenizers v1: encode, decode and scaling, measured

    Hugging Face has released `tokenizers` v1, a major performance update that achieves 3-30x faster encoding than v0.23 while maintaining identical output and API compatibility. Key optimizations include a SIMD-accelerated splitter, a thread-local word cache, and an allocation-free BPE merge loop, ensuring tokenization doesn't bottleneck ML workflows.

    Hugging Facehuggingface.co10 min
  3. Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models

    Feyospace‑v1 presents a data‑centric training pipeline for cyber‑security agents, combining five systems (Choulea, SkyReal, Hongzwang, PSBreakup, Kreator) to generate and verify 164 k long‑context trajectories across diverse exploit environments. The resulting checkpoints improve baseline performance by ~24% on CyberGym and achieve a 63% verified success rate, ranking top among similarly‑sized op…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

    This paper introduces label-free bias-only Test-Time Reinforcement Learning (TTRL), which uses majority-vote pseudolabels and optimizes only ~100K bias parameters. It achieves 76.67% accuracy on MATH-500 with Qwen2.5-7B, optimizing 76,000x fewer parameters than full-parameter TTRL, demonstrating substantial adaptation from a tiny subspace.

    Hugging Face Daily Papersarxiv.org1 minpaper