Skip to content
Open access

From data chaos to physically interpretable deterministic mapping

Jul 2026 · Nature Communications · Vol 17 · 0 citations · 59 references
Medicine

TL;DR

It is shown that the method consistently identifies compact governing equations while maintaining strong long-horizon predictive accuracy across canonical nonlinear systems and representative industrial processes, even under noisy and distribution-shifted data.

Abstract

Discovering governing equations directly from observational data remains a fundamental challenge in science and engineering, particularly when measurements are noisy, high-dimensional, or multi-scale. Existing approaches often cast equation discovery as a regression problem that selects candidate terms to fit observed trajectories, which can limit structural stability and identifiability under realistic data conditions. We propose a structured operator-learning framework that reformulates equation discovery as a constrained dynamical inference problem integrating spectral decomposition, physics-guided sparse projection, and cross-view consistency regularization within a unified architecture. By decomposing dynamics into scale-resolved components and enforcing invariance across perturbed observations, the framework promotes stable and interpretable equation recovery. Here, we show that the method consistently identifies compact governing equations while maintaining strong long-horizon predictive accuracy across canonical nonlinear systems and representative industrial processes, even under noisy and distribution-shifted data. Here, the authors propose structured operator learning with spectral decomposition, sparse regression, and cross-view regularization to recover stable, interpretable governing equations under noisy, high-dimensional, and distribution-shifted data.

Read PDF

Similar papers

Open access Jul 2026

Identifying nonlinear dynamical systems using subset regression.

The data-driven discovery of governing equations for dynamical systems has emerged as a transformative paradigm, enabling the extraction of interpretable and generalizable models from observational data. While modern techniques have advanced this field, traditional subset regression remains a foundational yet underutilized tool due to its reliance on uncorrelated residuals, a requirement often violated by time-series data. In this work, we revisit subset regression to identify dynamical systems governed by ordinary differential equations (ODEs), partial differential equations (PDEs), and differential algebraic equations (DAEs). We propose subset regression with known number of active features (sub-KNAFE), a user-determined sparsity mechanism that flexibly adapts to various complex nonlinear systems, while retaining the computational efficiency and inherent interpretability of traditional subset regression. We integrate sub-KNAFE with the SINDy framework, overcoming the limitation of subset regression in dynamical system identification. Numerical tests across a range of signal-to-noise ratios and dataset sizes demonstrate sub-KNAFE's superior noise robustness and data efficiency. Practical utility for sub-KNAFE is validated on two real-world datasets: the classic Lynx-Hare ecological population data and the ISO New England power system dataset, demonstrating its strong potential for practical deployment in scientific discovery and engineering applications.

Weizhen Li, Qiang Fu, Yifan Hong et al. · 0 citations
Review Open access May 2026

Data-Driven Identification of Stochastic Dynamical Systems

Identifying stochastic dynamical systems from observational data remains a major challenge in applied mathematics and engineering, particularly when complex systems are influenced by random perturbations and incomplete empirical information. This comprehensive review aims to examine state-of-the-art data-driven methods for discovering governing equations, estimating parameters, and predicting the behavior of stochastic dynamical systems. The review systematically analyzes key methodological approaches, including Sparse Identification of Nonlinear Dynamics (SINDy), Dynamic Mode Decomposition (DMD) and its extensions, Koopman operator theory, neural ordinary differential equations, and Bayesian inference. Each approach is evaluated in terms of its theoretical foundations, computational requirements, robustness to noise, and applicability to different classes of stochastic systems. Drawing on numerical experiments and real-world case studies, the findings show that no single method consistently outperforms others across all scenarios. Instead, hybrid approaches that integrate physics-informed constraints with machine learning demonstrate the strongest potential for advancing data-driven system identification. The review concludes that future research should address real-time identification, uncertainty quantification, and the integration of multi-fidelity data sources to improve the reliability and scalability of stochastic system modeling. This work contributes a comprehensive framework for guiding researchers and practitioners in selecting and implementing appropriate identification methods for stochastic dynamical systems.

Rishav Jha, Kameshwar Sahani, S. K. Sahani et al. · 0 citations
Preprint Jul 2026

Neural operator discovery from heterogeneous trajectories

Neural operators provide data-driven mappings for modeling dynamical systems. Extending them to families of systems typically requires explicit conditioning variables such as physical parameters, geometries, or boundary conditions. In many real-world settings, these quantities are unobserved. Here, we formulate neural operator discovery (NOD) as the problem of learning both shared solution operators and system-specific variation directly from heterogeneous trajectories without access to labeled governing factors. We introduce a factorized latent-conditioning formulation that jointly learns a neural operator and a low-dimensional latent representation through factorized prediction, trajectory-decoupled sampling, and dimension selection. Across diverse systems, the learned latent representation captures the intrinsic dimensionality of system variation and organizes system instances in a smooth and approximately invertible latent structure aligned with the underlying governing factors. This organization enables generalization to previously unseen system instances, including zero-shot extrapolation across regimes and stable long-horizon prediction. These results establish an interpretable paradigm for operator learning in the absence of explicit factor supervision.

Zituo Chen, Qiaofeng Li, Jiaxin Hu et al. · 1 citation
Preprint Jul 2026

Discovering Latent Response Laws in Forced Physical Systems

Governing equations provide compact descriptions of physical systems, yet the variables in which they are simple are often hidden in high-dimensional measurements. This challenge is sharper for forced systems, whose responses depend on both intrinsic dynamics and time-dependent inputs. Here we introduce FLARE, a forced latent autoencoder for response equations that learns compact response coordinates, identifies sparse input-dependent latent dynamics and decodes equation rollouts to full responses. By estimating latent dimension from data and separating state estimation from external forcing, FLARE enables forecasts to be initialized from past responses and driven by prescribed future inputs. Across known dynamical systems, application-scale forced responses and visual observations, FLARE recovers compact forced dynamics and predicts long-horizon high-dimensional responses under inputs not used for training. By turning learned coordinates into a dynamical interface, FLARE extends equation discovery to systems whose effective states are hidden within complex observations, providing a route for interpretable modelling and prediction of high-dimensional responses in forced dynamical systems.

Yi Zhu, Su Chen, Xiaojun Li et al. · 0 citations
Preprint Jul 2026

Learning Population-Level Dynamics through a Latent Fokker--Planck Model and Discrepancy Transport Maps

Many scientific and engineering systems are observed as time-indexed probability distributions whose governing dynamics are unknown and whose individual trajectories are unavailable. These settings challenge conventional system-identification approaches that rely on trajectory correspondence or prescribed evolution equations. This work presents a population-level inference framework that recovers latent stochastic dynamics directly from snapshot probability distributions by decomposing the observed evolution into an intrinsic latent stochastic process and a discrepancy transport map that captures geometric deformation between the latent and observed probability spaces. The latent dynamics are modeled using an Ornstein--Uhlenbeck process, providing a closed-form solution to the associated Fokker--Planck equation, while the discrepancy transport map is parameterized through the Knothe--Rosenblatt rearrangement with monotone neural networks. To mitigate the non-uniqueness inherent in the latent--transport decomposition, the transport map is regularized using a deformation energy motivated by hyperelasticity, promoting smooth, physically interpretable deformations while reducing unnecessary complexity. The latent stochastic model and discrepancy transport map are learned jointly through a unified optimization problem defined over probability distributions. Numerical examples involving nonlinear and multimodal distributional dynamics demonstrate that the proposed framework accurately reconstructs complex probability evolution while preserving a compact and analytically tractable latent representation. The proposed formulation provides a general framework for population-level dynamical inference and establishes a foundation for extending latent stochastic models and transport-based learning to more general and higher-dimensional systems.

Chengyang Huang, K. Garikipati · 0 citations
Preprint Aug 2026

Symbolic Neural ODEs: Learning interpretable models from time-series data

Multi-step training of sparse, interpretable models of dynamical systems directly from time-series data yields models with accurate short-term dynamics and strong agreement in long-time statistical properties, including mean, variance, and Lyapunov exponents.

N. Boddupalli, J. Moehlis · 0 citations