Skip to content

Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks

Jun 2026 · arXiv.org · Vol abs/2606.29519 · 0 citations
Computer Science Physics

TL;DR

It is shown that the asymptotic decay behavior of f is not fixed by the architecture and emerges from the coupling between the state dynamics and parameter dynamics, settling into either a collapsed regime (fast, exponential forgetting) or an extended, anti-collapsed regime (slow, power-law forgetting).

Abstract

Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data. This fade is captured by an envelope $f(\ell)$. An exponential fade makes the data needed to learn a lag-$\ell$ dependence grow exponentially, putting long horizons out of reach; a power-law fade keeps the cost polynomial. We show that the asymptotic decay behavior of $f(\ell)$ is not fixed by the architecture. Instead, it emerges from the coupling between the state dynamics and parameter dynamics, settling into either a collapsed regime (fast, exponential forgetting) or an extended, anti-collapsed regime (slow, power-law forgetting). The intuition is a competition within these coupled dynamics. Training drives the network's effective time scales toward short ones, while rare, heavy-tailed fluctuations of the learning dynamics push a few of them to very long values. Along the route studied here, the extended regime survives only when these heavy-tailed pushes are strong enough to balance the pull. We make this mathematically precise with a coarse-grained stochastic process and derive an explicit threshold at which this route to the extended regime becomes available. A single exponent, the spectral exponent~$\beta$, then governs both the spread of time scales and how slowly the network forgets. Realizing the regime in practice needs one more ingredient: the joint action of the architecture and the optimizer must be able to hold such a broad spread. A network whose capacity to generate broad time-scale spectra is severely constrained still collapses, even when supplied with strong heavy-tailed forcing. Heavy-tailed fluctuations thus act not as noise to be suppressed, but as the mechanism that sustains long-range learning.

View source

Similar papers

Preprint Jul 2026

Contraction versus Recurrence: An Exponential Separation in Observation-Based Prediction of Deterministic Dynamics

Given a scalar observable of an ergodic dynamical system with a low-dimensional attractor, two families of methods reconstruct and predict the underlying state: recurrence-based methods (the method of analogues and its descendants), which wait for the trajectory to return to an $\varepsilon$-neighborhood of a previously observed state, and observer-based methods, which fit a converging state estimator on the delay reconstruction. We formalize and empirically verify an exponential separation between the two: the expected cost of recurrence scales as $\varepsilon^{-d}$, where $d$ is the pointwise dimension of the invariant measure (a consequence of the Kac lemma and quantitative Poincare recurrence), whereas a detectable linear observer converges in $\Theta(\log(1/\varepsilon)/(1-\rho(A_{cl})^2))$ steps, where $\rho(A_{cl})$ is the closed-loop spectral radius of the Riccati fixed point. Both laws are verified numerically (return-time exponent $-1.8$ on the Lorenz attractor against the theoretical $-2.05$; observer cost linear in $\log(1/\varepsilon)$ with $R^2=1.000$ and in $(1-\rho^2)^{-1}$ with $R^2=0.985$), yielding a measured cost gap of $\sim 10^{9}$ at $\varepsilon=10^{-6}$ for $d\approx 2$. We complement the theorem with an admission protocol (the Kac-Riccati gate) deciding whether a signal lies inside the theorem's class, via surrogate-data prediction gating; it also explains the folklore of"universal"fractal dimensions as a dataset-size artifact bounded by $2\log_{10}N$. On real data the gate admits the Santa Fe laser benchmark ($\hat D_2=2.0$) and refuses the monthly sunspot series, reproducing the settled resolution of historical low-dimensionality claims. All results reproduce from a single verification script (17/17 checks).

Pavel Popovich · 0 citations
Preprint Aug 2026

Differential-Embedding Reconstruction of Dynamical Systems from Scalar Time Series

We study the reconstruction of an unknown dynamical system from a single noisy scalar time series. The goal is to recover the underlying dynamics for forecasting. We introduce a method that uses differential embedding coordinates to identify a rational closure of the embedding dynamics directly from data. The closure is identified through a weak-form regression pipeline, which avoids unstable pointwise differentiation of noisy data. When applied to noise-free Lorenz and R\"ossler systems, the method recovers closures that support long forecasts across a broad ensemble of realizations ($18.1$ and $7.1$ Lyapunov times respectively). Under $15$--$30\%$ additive Gaussian noise, performance becomes system-dependent. For the Lorenz system, forecast horizons remain short even in the best cases, whereas the R\"ossler system generally performs better in absolute terms, though not once normalized by the Lyapunov time. Our proposed method recovers directly interpretable closure coefficients which we compared against the known analytic closures of the Lorenz and R\"ossler systems.

A. Shaa, C. Guet · 0 citations
Open access Jul 2026

Resolution-Induced Collapse in Quantized Nonlinear Dynamics: A Finite-Horizon Structural Framework

Finite-precision implementation fundamentally changes nonlinear dynamical systems by replacing continuous-state evolution with deterministic dynamics on a finite set of representable states. This study examines when that change becomes structurally important over a finite observation horizon. Quantization is treated as a resolution constraint, and an operational separation scale δsep(T0,T;ε) is introduced to compare the implementation resolution with the attractor detail exposed by a reference trajectory. The ratio η=Δ/δsep and its associated critical bit width bc(T) are used as protocol-dependent measures for precision screening. Experiments on the Hénon map, a Lorenz system integrated by fixed-step fourth-order Runge–Kutta, and the Logistic map show strong system dependence. Hénon exhibits broad, non-monotonic finite-state reshaping across bit width, whereas Lorenz remains in a low-complexity regime over a wider low-bit range before recovering more complex recurrent behavior. Results from 100 selected occupied quantized attractor positions show that entropy alone is insufficient to characterize collapse; recurrence and transient lengths provide complementary information about orbit organization. A 21-horizon Lorenz study produces stepwise changes and a long plateau in bc(T), rather than a smooth linear scaling law. For Hénon, the largest-horizon crossing is resolved at bc=25.457, corresponding to a minimum integer bit width of 26 under the declared estimator. An alternative estimator gives materially different crossing values, while the Logistic map also shows strongly non-monotonic behavior. Overall, η and bc(T) provide useful implementation-oriented screening measures, but they are not universal thresholds or hardware guarantees.

Lei Zhang · 0 citations
Preprint Jul 2026

Extrapolating the emergence of Hamiltonian chaos with random-feature Hamiltonian neural networks

Machine learning of Hamiltonian dynamics has driven growing interest in Hamiltonian neural networks (HNNs), which encode Hamilton's equations of motion into the learning architecture. Despite this progress, it remains unknown whether such networks can predict dynamical regimes absent from their training data, in particular the broad chaotic sea that emerges beyond the observed parameter interval. We address this question using a parameter-aware random-feature Hamiltonian neural network (RF-HNN). Trained using data from only a small number of control-parameter values at which invariant tori dominate, the RF-HNN predicts autonomous long-time dynamics at unseen parameter values where mixed phase space develops and chaotic regions expand, with no data from that regime used in training or model selection. The method is demonstrated across four two-degree-of-freedom Hamiltonian families, including the H\'enon-Heiles system. Using Poincar\'e-section geometry and finite-time Lyapunov exponents, we show that the RF-HNN reproduces the breakup of regular structures and the emergence and growth of chaotic regions, whereas conventionally trained HNNs with the same Hamiltonian structure remain too regular. These results show that what decides parameter extrapolation is not Hamiltonian structure alone but how the fitted Hamiltonian continues in the control parameter. To our knowledge, this is the first demonstration that a learned Hamiltonian can qualitatively extrapolate from predominantly regular dynamics into a broad chaotic sea absent from training.

Jaesung Choi · 0 citations
Preprint Jul 2026

Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries

Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight decay, together with a locally quadratic extension to nonlinear neural networks. Our analysis reveals a distinguished population-active component of the empirical null space, which we call the grokking subspace. Along this subspace, the training predictions remain unchanged, leaving weight decay as the sole restoring force and giving rise to a slow dissipative relaxation governed by an exact discrete-time and continuous-time law. We show that only this subspace contributes to the slow asymptotic decay of the population risk and derive explicit iteration-scale predictions for the grokking time, recovering the familiar $(1-\beta)/(\eta\lambda)$ scaling in the weak-regularization regime. The theory further predicts distinct effects of optimizer choice, distinguishing coupled $L_2$ regularization from decoupled weight decay, and yields causal predictions for interventions that modify the grokking component. We verify all theoretical identities without fitted parameters in a synthetic model where every subspace and relaxation rate is computable in closed form. We further observe genuine delayed generalization in modular addition, where the measured delay follows the predicted scaling and the late-time relaxation agrees closely with the theoretical clock.

Taeyoung Kim · 0 citations
Open access Jul 2026

Structured Fluctuations and the Information Dynamics of Self-Maintenance in Growing Neural Cellular Automata

Growing Neural Cellular Automata (GNCA) are capable of robust self-maintenance and self-repair, yet the internal dynamical mechanisms that support these capabilities remain poorly understood. Here, we investigate the role of internal fluctuations—temporal micro-variability of hidden channel states—in a trained GNCA model, hypothesizing that they constitute a functional component of the dynamics rather than merely residual stochastic noise. We analyzed the trained model through dynamical-systems analysis (low-dimensional embedding and recurrence analysis of collective state trajectories) and information-theoretic analysis (transfer entropy and partial information decomposition), including its response to localized damage and to suppression of small-magnitude updates. These analyses show that internal fluctuations are spatially structured, dynamically coupled to an attracting collective state, and associated with distributed small-magnitude updates that contribute to damage recovery. Damage induces a global deviation in latent state space followed by gradual re-convergence, and suppressing distributed small-magnitude updates associated with baseline fluctuation dynamics outside a permissive radius that encompasses the majority of the cells significantly impairs recovery. Transfer entropy analysis characterizes a spatially differentiated repair response: corrective inward flow near the damage site coexists with outward perturbation propagation at greater distances. Partial information decomposition further suggests a regime shift from synergy-dominant resting computation to redundancy-increased coordination during recovery. These findings indicate that GNCA self-maintenance and self-repair emerge from high-dimensional nonlinear collective dynamics in which internal fluctuations serve as a functional component supporting information flow, coordination, and return toward an attracting recurrent state.

A. Masumori, Hiroki Sato, Takashi Ikegami · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.