Adam
Adam (Adaptive Moment Estimation) is an optimisation algorithm for training neural networks, introduced by Diederik Kingma and Jimmy Ba in 2014. It's an extension of Stochastic Gradient Descent.
Adam keeps an exponential moving average of the gradients (the first moment) and of the squared gradients (the second moment) for each parameter. It uses these to give each parameter its own effective step size, and applies a bias correction because both averages start at zero.
The commonly used default hyperparameters are , and . It works well out of the box on a lot of problems, which is why it's one of the most popular optimisers in deep learning.