Hugging Face Daily PapersYanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi1 min readpaperadvanced
Scaling Laws for Looped Mixture of Experts
Summary
This paper introduces Loop Scaling Laws, which jointly model recurrence and sparsity in Mixture-of-Experts (MoE) transformers. It finds that recurrence and sparsity offer complementary efficiency gains, improving prediction accuracy and enabling more efficient large model design.
- New Loop Scaling Laws model recurrence and MoE sparsity together for transformer performance.
- Sparsity provides ~3x active-parameter efficiency in MoE models.
- Recurrence yields ~2x total-parameter efficiency, especially for reasoning tasks.
- Jointly scaling recurrence and sparsity further improves performance frontiers.
This work provides a principled framework for researchers and engineers to design more efficient and performant large language models by leveraging both recurrence and sparse MoE architectures.
8/10