Hall of FameDiederik P. Kingma, Jimmy Ba201432 min readpaperintermediate
Adam: A Method for Stochastic Optimization
Summary
This paper introduces Adam, a first-order gradient-based optimization algorithm that adaptively estimates first and second moments of gradients. It computes individual learning rates for different parameters, making it efficient and robust for large-scale, high-dimensional machine learning problems with noisy or sparse gradients.
- Adam computes adaptive learning rates for each parameter using exponential moving averages of gradients and squared gradients.
- It includes bias correction for the initial moment estimates, which is crucial for performance during early training steps.
- The algorithm combines advantages of AdaGrad (sparse gradients) and RMSProp (non-stationary objectives).
- Parameter updates are invariant to gradient rescaling, and effective stepsizes are approximately bounded by the learning rate hyperparameter.
This paper introduced Adam, which became one of the most popular and effective optimizers for deep learning, significantly simplifying and improving the training of complex neural networks.
9/10