proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameTom B. Brown et al.2020166 min readpaperadvanced

Language Models are Few-Shot Learners

Summary

This paper introduces GPT-3, a 175-billion-parameter autoregressive language model, demonstrating that scaling model size significantly improves few-shot learning. It achieves strong performance on many NLP tasks by conditioning on text instructions and examples, often without needing gradient updates or fine-tuning.

  • GPT-3 is a 175B parameter autoregressive language model, 10x larger than any previous non-sparse model at the time.
  • It performs "in-context learning" by taking task instructions and demonstrations directly in the input text, without gradient updates.
  • Few-shot performance scales significantly with model size, often matching or exceeding prior state-of-the-art fine-tuning approaches.
  • GPT-3 shows strong results on diverse tasks including translation, question-answering, arithmetic, and human-like text generation.

This paper was foundational, demonstrating that sufficiently large language models can exhibit strong few-shot learning capabilities without task-specific fine-tuning, fundamentally shifting the paradigm for NLP model development and application.

9/10

Related reading

  1. Scaling Laws for Neural Language Models

    This paper empirically studies scaling laws for neural language model performance, finding that cross-entropy loss scales as a power-law with model size, dataset size, and compute. It shows that optimal compute-efficient training involves using very large models, training on relatively modest data, and stopping significantly before convergence.

    Hall of Famearxiv.org67 minpaper
  2. Training Compute-Optimal Large Language Models

    The paper derives a compute‑optimal scaling law showing model size and training tokens should grow together, and validates it by training a 70B‑parameter model (Chinchilla) on 1.4 T tokens that outperforms much larger LLMs.

    Hall of Famearxiv.org66 minpaper
  3. Convergent Emergence of In-Context Learning Across Modalities

    The paper proposes the Convergent Emergence Hypothesis that few‑shot in‑context learning (ICL) shares a common difficulty profile across domains. Using a unified task suite, the authors evaluate ICL on six modalities—language, genome, integer sequences, time‑series, images, and proteins—showing that paired‑mapping ICL emerges in all and that per‑task benefits correlate across five modalities, sup…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. EasyPPO: Stabilizing the Critic Is Key

    EasyPPO addresses two key instability issues in PPO's critic when training large language models: biased filtering of truncated rollouts and heterogeneous return noise. It introduces actor-only overlong filtering, noise-normalized critic regression, and smaller critic mini-batches, achieving significant performance gains over vanilla PPO across various tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Draft-KV: Learning Useful Latent Communication Between Language Models

    Draft‑KV introduces a lightweight interface that passes a sharer LLM’s key‑value cache to a frozen receiver via gated attention, training only 1.05 M parameters. This latent communication lifts a 0.5 B receiver from ~37 % to 78 % on MMLU‑Redux and scales with the sharer size, while the sharer can be omitted at inference.

    Hugging Face Daily Papersarxiv.org1 minpaper