Hall of FameTom B. Brown et al.2020166 min readpaperadvanced
Language Models are Few-Shot Learners
Summary
This paper introduces GPT-3, a 175-billion-parameter autoregressive language model, demonstrating that scaling model size significantly improves few-shot learning. It achieves strong performance on many NLP tasks by conditioning on text instructions and examples, often without needing gradient updates or fine-tuning.
- GPT-3 is a 175B parameter autoregressive language model, 10x larger than any previous non-sparse model at the time.
- It performs "in-context learning" by taking task instructions and demonstrations directly in the input text, without gradient updates.
- Few-shot performance scales significantly with model size, often matching or exceeding prior state-of-the-art fine-tuning approaches.
- GPT-3 shows strong results on diverse tasks including translation, question-answering, arithmetic, and human-like text generation.
This paper was foundational, demonstrating that sufficiently large language models can exhibit strong few-shot learning capabilities without task-specific fine-tuning, fundamentally shifting the paradigm for NLP model development and application.
9/10