proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameJared Kaplan et al.202067 min readpaperadvanced

Scaling Laws for Neural Language Models

Summary

This paper empirically studies scaling laws for neural language model performance, finding that cross-entropy loss scales as a power-law with model size, dataset size, and compute. It shows that optimal compute-efficient training involves using very large models, training on relatively modest data, and stopping significantly before convergence.

  • LM performance follows precise power-laws with model size (N), dataset size (D), and compute (C).
  • Architectural details like width/depth have minimal impact on performance within a wide range.
  • Larger models are significantly more sample-efficient, reaching the same performance with fewer data points.
  • Optimal compute allocation prioritizes training very large models and stopping early, rather than training small models to convergence.

This paper established the foundational empirical scaling laws that guide the design and training of modern large language models, informing resource allocation and architectural choices for practitioners.

9/10

Related reading

  1. Training Compute-Optimal Large Language Models

    The paper derives a compute‑optimal scaling law showing model size and training tokens should grow together, and validates it by training a 70B‑parameter model (Chinchilla) on 1.4 T tokens that outperforms much larger LLMs.

    Hall of Famearxiv.org66 minpaper
  2. Language Models are Few-Shot Learners

    This paper introduces GPT-3, a 175-billion-parameter autoregressive language model, demonstrating that scaling model size significantly improves few-shot learning. It achieves strong performance on many NLP tasks by conditioning on text instructions and examples, often without needing gradient updates or fine-tuning.

    Hall of Famearxiv.org166 minpaper
  3. Efficient Estimation of Word Representations in Vector Space

    This paper introduces two novel log-linear model architectures for efficiently computing continuous word vector representations from very large datasets. These models achieve state-of-the-art accuracy on syntactic and semantic word similarity tasks with significantly lower computational cost than previous neural network approaches.

    Hall of Famearxiv.org27 minpaper
  4. Scaling Laws for Looped Mixture of Experts

    This paper introduces Loop Scaling Laws, which jointly model recurrence and sparsity in Mixture-of-Experts (MoE) transformers. It finds that recurrence and sparsity offer complementary efficiency gains, improving prediction accuracy and enabling more efficient large model design.

    Hugging Face Daily Papersarxiv.org1 minpaperHN2