Skip to content

The Role of Gradient Modification in Heavy-Tailed Nonconvex Stochastic Min-Max Optimization

Sep 2026 · 0 citations
Mathematics Computer Science

TL;DR

It is shown that vanilla SGDA, without any modification to its update rule, can converge under heavy-tailed noise in both nonconvex-strongly-concave (NC-SC) and nonconvex-concave (NC-C) settings, establishing the first convergence guarantees for SGDA in these regimes.

Abstract

Stochastic min-max optimization has attracted increasing attention due to its applications in modern machine learning, while existing theoretical studies mainly rely on the bounded variance assumption for stochastic gradients. Under heavy-tailed noise, where stochastic gradients only possess a finite $p$-th moment for $p\in(1,2]$, gradient clipping or normalization is commonly believed to be necessary to guarantee convergence. In this work, we revisit stochastic min-max optimization under heavy-tailed noise and provide a comprehensive theoretical study of stochastic gradient descent ascent (SGDA). We first show that vanilla SGDA, without any modification to its update rule, can converge under heavy-tailed noise in both nonconvex-strongly-concave (NC-SC) and nonconvex-concave (NC-C) settings, establishing the first convergence guarantees for SGDA in these regimes. Beyond unregularized problems, we further investigate regularized stochastic min-max optimization, where directly incorporating gradient normalization into proximal updates is nontrivial due to the incompatibility between normalization and proximal structures. We overcome this difficulty by developing new clipping-free algorithms, i.e., Stoc-TRGDAM and Stoc-TRGDmax, and they both can achieve the optimal dependence on the target accuracy without using gradient clipping.

View source

Similar papers

#machine learning Preprint Sep 2026

High-Probability Convergence of SGD via Batched Updates

Stochastic gradient descent (SGD) is the primary workhorse for large-scale optimization. While the average behavior of its iterates, typically characterized by mean-squared error bounds, is well-understood, obtaining high-probability guarantees for the last iterate remains challenging. Prior approaches to this problem...

Feng Zhu, Robert W. Heath, Aritra Mitra · 0 citations
#machine learning Preprint Sep 2026

Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise and H\"{o}lder Smoothness

Classical convergence guarantees for stochastic gradient methods typically assume Lipschitz-smooth objectives and finite-variance gradient noise, both frequently violated in practice. In contrast, we study nonconvex stochastic optimization under the joint relaxation of these assumptions: objectives with $(L,s)$-H\"olde...

M. Zaman, Anirbit Mukherjee · 0 citations
#machine learning Preprint Sep 2026

TR-SSQP: A Trust-Region Method for Constrained Stochastic Optimization under Heavy-Tailed Noise

We consider stochastic nonlinear optimization problems with deterministic equality constraints. While unconstrained stochastic optimization is well understood, the interplay between optimality and feasibility in the constrained setting poses significant challenges. Moreover, existing theoretical guarantees for constrai...

Hao-Xuan Wang, Yu-Chen Fang, Sen Na · 0 citations
Preprint Sep 2026

Gradient-Free Methods for Stochastic Convex Optimization with Stochastic Functional Constraints

We develop accelerated gradient-free methods for stochastic convex optimization with constraints defined by expectations. Our batched primal-dual sliding method uses two-point evaluations sharing a random sample and guarantees expected objective error and expected maximum constraint violation at most $\varepsilon$. It...

Vadim Abronin, A. Gasnikov, D. Dvinskikh · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.