Understanding the Optimization Dynamics of Large Language Models: A Perspective from Non-Equilibrium Statistical Physics
Abstract
The training process of modern deep learning models, particularly Large Language Models (LLMs), exhibits complex stochastic behaviors that are difficult to explain using traditional optimization theories. In this paper, we propose a novel perspective by modeling the optimization process of neural networks as a non-equilibrium physical process driven by generalized Langevin dynamics. We establish a rigorous mapping between neural network optimization variables and thermodynamic quantities. Furthermore, we derive and empirically validate the generalized Fluctuation-Dissipation Theorem (FDT) within the framework of the Adam optimizer. Through spectral analysis, we demonstrate a remarkably high linear correlation between the gradient noise power spectrum and the dissipative memory kernel, confirming an intrinsic thermodynamic balance. Additionally, through gradient power spectrum analysis during LLM fine-tuning, we discover the spontaneous emergence of Self-Organized Criticality (SOC), characterized by a distinct $1 / f$ noise signature indicating long-range temporal correlations. Our findings suggest that maintaining the system in this critical state acts as a universal optimization attractor that maximizes information processing capacity, providing an optimal balance between training stability and exploration efficiency.