proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi1 min readpaperadvanced

Scaling Laws for Looped Mixture of Experts

Summary

This paper introduces Loop Scaling Laws, which jointly model recurrence and sparsity in Mixture-of-Experts (MoE) transformers. It finds that recurrence and sparsity offer complementary efficiency gains, improving prediction accuracy and enabling more efficient large model design.

  • New Loop Scaling Laws model recurrence and MoE sparsity together for transformer performance.
  • Sparsity provides ~3x active-parameter efficiency in MoE models.
  • Recurrence yields ~2x total-parameter efficiency, especially for reasoning tasks.
  • Jointly scaling recurrence and sparsity further improves performance frontiers.

This work provides a principled framework for researchers and engineers to design more efficient and performant large language models by leveraging both recurrence and sparse MoE architectures.

8/10

Related reading

  1. IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

    IntBMoE introduces block‑level conditioning to MoE, decoupling token participation, compute execution, and memory materialization. A hypernetwork merges all experts into a composed expert per block, while routing remains sparse. Dual‑Path Residual Gating further mixes two composed paths. Experiments show consistent gains on vision, language, and recommendation tasks, and the model is live in AMap…

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. Scaling Laws for Neural Language Models

    This paper empirically studies scaling laws for neural language model performance, finding that cross-entropy loss scales as a power-law with model size, dataset size, and compute. It shows that optimal compute-efficient training involves using very large models, training on relatively modest data, and stopping significantly before convergence.

    Hall of Famearxiv.org67 minpaper
  3. WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

    WhiteMatter introduces all-to-all cross-layer connections in Transformers by mixing past-token representations from any depth into shared KV cache channels. This approach allows for performance comparable to 50% larger standard Transformers or improved performance with half the KV cache size. It also addresses training slowdowns with a novel cyclic iteration method.

    Hugging Face Daily Papersarxiv.org1 minpaper