Skip to content

Author

Davide Gallon

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Strong error analysis for the stochastic momentum optimizer

Stochastic gradient descent (SGD) optimization schemes are the methods of choice for the optimization of deep neural networks (DNNs) in artificial intelligence (AI) systems. Often not the standard SGD method is used but instead suitable accelerated, adaptive, and/or normalized variants of standard SGD such as Adam, AdamW, and MUON are employed to train large scale AI systems in practically relevant settings. The acceleration (higher order convergence speed) in all these popular optimizers relies on the momentum SGD optimizer. In this work we provide a rigorous error analysis for the momentum SGD optimizer. In particular, we establish convergence rates for the momentum optimizer in terms of the size of the learning rate (step size), the size of the mini-batch, and the size of the one-point convexity constant.

Davide Gallon, Arnulf Jentzen · 0 citations