TAI BUI
← Glossary
Glossary

What is AdamW?

Adam with decoupled weight decay. In standard Adam, L2 regularization gets scaled by the adaptive learning rate per parameter, which is not what you want. AdamW applies weight decay directly to the weights, independent of the gradient statistics. The default optimizer for training transformers.

What people say

Adam but better

Why it's called that