proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameJacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova201843 min readpaperadvanced

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Summary

BERT introduced Bidirectional Encoder Representations from Transformers, pre-trained using a Masked Language Model and Next Sentence Prediction. This enabled deep bidirectional context, achieving new state-of-the-art results across eleven NLP tasks with minimal fine-tuning.

  • BERT uses a deep bidirectional Transformer encoder, unlike prior unidirectional models.
  • Pre-training involves a Masked Language Model (MLM) to learn from full left and right context.
  • A Next Sentence Prediction (NSP) task is used to pre-train text-pair representations.
  • It achieved state-of-the-art performance on 11 NLP tasks by fine-tuning with minimal task-specific layers.

This paper introduced a foundational architecture and pre-training strategy that revolutionized NLP, enabling significant performance gains and widespread adoption of large pre-trained models.

9/10

Related reading

  1. Efficient Estimation of Word Representations in Vector Space

    This paper introduces two novel log-linear model architectures for efficiently computing continuous word vector representations from very large datasets. These models achieve state-of-the-art accuracy on syntactic and semantic word similarity tasks with significantly lower computational cost than previous neural network approaches.

    Hall of Famearxiv.org27 minpaper
  2. Language Models are Few-Shot Learners

    This paper introduces GPT-3, a 175-billion-parameter autoregressive language model, demonstrating that scaling model size significantly improves few-shot learning. It achieves strong performance on many NLP tasks by conditioning on text instructions and examples, often without needing gradient updates or fine-tuning.

    Hall of Famearxiv.org166 minpaper
  3. Transformers Explained Visually

    The article walks through the core components of a text‑generative Transformer—embedding, multi‑head self‑attention, MLP, and output projection—using GPT‑2 small as a concrete example. It shows the dimensions, parameter counts, and step‑by‑step calculations that underlie token prediction.

    Hacker News front pagegithub.io11 minHN63388
  4. Attention Is All You Need

    The paper proposes the Transformer, a sequence‑to‑sequence model that relies solely on self‑attention, eliminating recurrence and convolutions. It achieves state‑of‑the‑art translation BLEU scores while training orders of magnitude faster.

    Hall of Famearxiv.org27 minpaper