Skip to content

HyperTransfer: Understanding the Equivalence between Base Optimizer and Hyperball

Sep 2026 · 0 citations · 26 references
Computer Science

TL;DR

HyperTransfer is proposed, which constructs a Hyperball optimizer that reproduces the dynamics of a target Base Optimizer using only its initialization and learning-rate schedule, without running the target optimizer itself, to extend the framework to non-scale-invariant networks.

Abstract

Hyperball optimizers constrain parameter norms and update only their directions, establishing a distinct paradigm for neural network optimization. Although this geometry appears fundamentally different from that of conventional Base Optimizers, which update both parameter norms and directions, we show that the two paradigms are dynamically equivalent for scale-invariant networks. Building on this equivalence, we propose HyperTransfer, which constructs a Hyperball optimizer that reproduces the dynamics of a target Base Optimizer using only its initialization and learning-rate schedule, without running the target optimizer itself. We further derive the inverse mapping and extend the framework to non-scale-invariant networks. Experiments show that both HyperTransfer and the inverse mapping produce loss trajectories nearly identical to those of their targets, suggesting that Hyperball dynamics are governed primarily by the induced effective learning-rate schedule and optimizer state.

View source

Similar papers

Preprint Aug 2026

RODE: A Radial-Orthogonal Decoupled Engine for Optimization

This work introduces RODE, which gives the radial and directional components separate update rules and step sizes in the matrix Frobenius norm, and suggests that decoupling radial and directional dynamics offers a more effective and controllable approach to matrix optimization.

Guo-Xiang Xu, Bince Qu, Qi Sun et al. · 0 citations
Preprint Aug 2026

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

Results establish ELR as a common coordinate linking LR scheduling, norm control, and loss dynamics, and Controlled interventions further show that weight decay and Hyperball shape loss dynamics primarily through the ELR schedules they induce.

Zi-Han Liu, Rui-Heng Zheng, Shaobo Zhang et al. · 3 citations
#artificial intelligence Preprint Sep 2026

Relative Generalization Invariance of LLM Pretraining

Large Language Model (LLM) pretraining performance is jointly shaped by three components of the training triplet: the optimizer, model architecture, and training data stream. However, how these components influence performance in distinct ways remains unclear. We take a first step toward isolating their effects by stud...

Feng-Zhuo Zhang, Shu-Che Wang, Sheng-Gui Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Equivariance Breaks the Learning Rate

Block normalization generally improves Adam and closes part of its gap to Muon across molecular force prediction on the rMD17 and MD22 datasets, QM9 molecular property prediction, and charged particle dynamics.

Andrei Manolache, Mathias Niepert · 0 citations
#artificial intelligence Review Aug 2026

Blog: Survey of Optimizers

This survey organizes recent optimizers and training optimization methods along four largely independent axes: temporal estimation, update geometry, horizon management, and representation and systems.

Ruo-Ran Xu · 0 citations
#machine learning Preprint Sep 2026

Muon Sublates the Edge of Stability in LLM Pretraining

Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability meet at a single learning-rate-dependent edge. We show that Mu...

Yan-Zhe Chen, Qifang Zhao, Xiao-Xiao Xu et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.