Skip to content

Equivariance Breaks the Learning Rate

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

Block normalization generally improves Adam and closes part of its gap to Muon across molecular force prediction on the rMD17 and MD22 datasets, QM9 molecular property prediction, and charged particle dynamics.

Abstract

Equivariant networks are commonly trained with Adam, yet recent work reports that matrix structured optimizers such as Muon can perform better, with the reasons for these gains only partly understood. We identify one source of this difference inside equivariant layers. An equivariant layer learns one channel mixing matrix $W_l$ per degree $l$, which we call an irrep block, and shares it across the $2l+1$ components, giving the expanded map $W_l \otimes I_{2l+1}$. This sharing sums gradient contributions across components and can produce different update scales under SGD. Adam's entrywise normalization reduces sensitivity to gradient scale, but neither optimizer directly controls the effective step size of each block. A single learning rate can therefore produce different effective step sizes across blocks. Muon instead controls the effective step size by approximately equalizing the singular values of each momentum matrix. We normalize each irrep block update by a single scalar, preserving its singular value ratios while letting the learning rate control its size. We implement this with spectral normalization or a simpler root-mean-square normalization. We evaluate spectral normalization in a controlled $\mathrm{SO}(3)$-equivariant model with a matched non-equivariant model. In this setting, the step size mismatch grows with width in the equivariant model but not in the non-equivariant model. We evaluate both variants across molecular force prediction on the rMD17 and MD22 datasets, QM9 molecular property prediction, and charged particle dynamics. Across these applications, block normalization generally improves Adam and closes part of its gap to Muon. These results highlight an overlooked interaction between equivariant architectures and their optimizers. Studying and designing the two together may help explain and address training difficulties often attributed to equivariance itself.

View source

Similar papers

#machine learning Preprint Sep 2026

Learning the identity: a case study of how SGD selects among functional decompositions

One might think that learning the identity function with a deep linear residual network is trivial - the path along residual connections already implements the identity, and so the network need only drive its weights to zero. However, this zero-weight solution is just one point on an entire manifold of population-loss...

Andy Arditi, Wei-An Xie, D. Bau et al. · 0 citations
Preprint Aug 2026

Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive

Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance, typically diagonal Fisher values. Per-parameter looks more flexible than per-layer, but each layer's diagonal Fisher is a weak summary of its actual curvature, missing the top-eig...

Brian B. Moser, Ahmed Anwar, T. Nauen et al. · 0 citations
Preprint Aug 2026

NAE: Normalizing AutoEncoder

This work proposes Normalizing Autoencoder (NAE), which employs a novel conditional loss that aligns the surrogate loss gradient with that of reconstruction loss, directly improving upon the current standard.

Muhammad Abdur Rafae, Niels Landwehr · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond the Matrix Sign: Quadratic Spectral Descent

Muon emerges as a strong competitor of the AdamW for LLM pretraining, because the matrix-wise update it employs can potentially incur smaller second-order penalty than the once dominating AdamW, which performs coordinate-wise update. However, the spectral flattening procedure in Muon is quite debatable since it discard...

Qiao-Zhe Zhang, Jun Sun, Ying-Zhuang Liu · 0 citations
#machine learning Preprint Sep 2026

Does a Shared Temperature Imply a Shared Angular Scale in Probabilistic Contrastive Learning?

In probabilistic contrastive learning, a shared temperature is commonly interpreted as a shared similarity scale, but this interpretation does not hold for high-dimensional distributional class representations. We study the exact von Mises-Fisher (vMF) probabilistic score used by ProCo when representation dimension and...

Ning-Kang Peng, Qian-Feng Yu, Jing Mao et al. · 0 citations
Preprint Aug 2026

The Sparsity Whisperer

A family of difference-informed pruning methods built upon this principle are introduced, suggesting that preserving output differences is a broadly useful and composable signal for post-training LLM sparsification.

Linghao Kong, Inimai Subramanian, Micah Adler et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.