Home /permanent

Adam

Adam (Adaptive Moment Estimation) is an optimisation algorithm for training neural networks, introduced by Diederik Kingma and Jimmy Ba in 2014. It's an extension of Stochastic Gradient Descent.

Adam keeps an exponential moving average of the gradients (the first moment) and of the squared gradients (the second moment) for each parameter. It uses these to give each parameter its own effective step size, and applies a bias correction because both averages start at zero.

The commonly used default hyperparameters are β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999 and ϵ=10−8\epsilon = 10^{-8}. It works well out of the box on a lot of problems, which is why it's one of the most popular optimisers in deep learning.