proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameDiederik P. Kingma, Jimmy Ba201432 min readpaperintermediate

Adam: A Method for Stochastic Optimization

Summary

This paper introduces Adam, a first-order gradient-based optimization algorithm that adaptively estimates first and second moments of gradients. It computes individual learning rates for different parameters, making it efficient and robust for large-scale, high-dimensional machine learning problems with noisy or sparse gradients.

  • Adam computes adaptive learning rates for each parameter using exponential moving averages of gradients and squared gradients.
  • It includes bias correction for the initial moment estimates, which is crucial for performance during early training steps.
  • The algorithm combines advantages of AdaGrad (sparse gradients) and RMSProp (non-stationary objectives).
  • Parameter updates are invariant to gradient rescaling, and effective stepsizes are approximately bounded by the learning rate hyperparameter.

This paper introduced Adam, which became one of the most popular and effective optimizers for deep learning, significantly simplifying and improving the training of complex neural networks.

9/10

Related reading

  1. Deep Residual Learning for Image Recognition

    The paper proposes reformulating deep layers as residual functions with identity shortcut connections, making it easy to train networks far deeper than before. Using this design, a 152‑layer ResNet achieved 3.57% top‑5 error on ImageNet, winning ILSVRC 2015.

    Hall of Famearxiv.org42 minpaper
  2. Generative Adversarial Nets

    This paper introduces Generative Adversarial Nets (GANs), a novel framework for training generative models. It pits a generator (G) against a discriminator (D) in a minimax game, where G tries to produce data that D cannot distinguish from real data, and D tries to correctly classify real vs. generated samples.

    Hall of Famearxiv.org20 minpaper
  3. Efficient Estimation of Word Representations in Vector Space

    This paper introduces two novel log-linear model architectures for efficiently computing continuous word vector representations from very large datasets. These models achieve state-of-the-art accuracy on syntactic and semantic word similarity tasks with significantly lower computational cost than previous neural network approaches.

    Hall of Famearxiv.org27 minpaper