Hall of FameJared Kaplan et al.202067 min readpaperadvanced
Scaling Laws for Neural Language Models
Summary
This paper empirically studies scaling laws for neural language model performance, finding that cross-entropy loss scales as a power-law with model size, dataset size, and compute. It shows that optimal compute-efficient training involves using very large models, training on relatively modest data, and stopping significantly before convergence.
- LM performance follows precise power-laws with model size (N), dataset size (D), and compute (C).
- Architectural details like width/depth have minimal impact on performance within a wide range.
- Larger models are significantly more sample-efficient, reaching the same performance with fewer data points.
- Optimal compute allocation prioritizes training very large models and stopping early, rather than training small models to convergence.
This paper established the foundational empirical scaling laws that guide the design and training of modern large language models, informing resource allocation and architectural choices for practitioners.
9/10