Skip to content
Open access

Causality beyond the Reichenbach principle: the degrees-of-freedom method

Aug 2026 · Scientific Reports · 0 citations

Abstract

Causal relationships underpin scientific reasoning across disciplines, yet reliable inference from observational data remains challenging. In many fields, from climate science to neuroscience, direct interventions are infeasible, and distinguishing genuine causality from correlations induced by hidden drivers is notoriously difficult. Classical approaches, from Granger causality to topological embedding methods, can be theoretically well founded for either stochastic or deterministic systems, but they usually do not cover both settings in the same framework, and only a few include an explicit hidden-driver class. In an earlier paper, we introduced a general framework for causal inference in dynamical systems based on the effective degrees of freedom of a system (df). The framework applies uniformly to deterministic and stochastic systems, in both discrete and continuous time. In the present paper, we develop an operational implementation of this framework based on conditional-entropy features and supervised classification, evaluate it on synthetic systems, and apply it to selected real time series as proof-of-applicability examples. Causal model classes are constructed using simulated coupled 0–1 Markov chains, and observed trajectories are mapped to binary processes and classified according to their df relationships. This procedure is designed to distinguish canonical causal classes, including directed driving, common cause, bidirectional coupling, and independence. The approach achieves high accuracy on controlled synthetic data, while performance under distribution shift depends on sample length and noise. In a small Hénon benchmark against Sauer–Sugihara, Granger causality, and transfer entropy, its advantage is not uniform, but is most visible in the full five-class task at longer sample length, where the common-driver class is retained. Applications to climate, ecological, and neural datasets demonstrate the breadth of the approach. In the climate example, different preprocessing and lag choices produced different class-score patterns, including a contemporaneous configuration and lagged configurations. We interpret these empirical examples as demonstrations of applicability rather than as definitive domain-specific causal conclusions. These findings suggest that df analysis can provide a broadly applicable and conceptually transparent route to disentangling causality from correlation in complex systems, provided that training design, preprocessing, and robustness checks are handled carefully.

Read PDF

Similar papers

#machine learning Preprint Aug 2026

Causal Local States: Scalable Simultaneous Causal Network Inference and Forecasting for Dynamical Systems

Reconstruction of the underlying networks with high fidelity and forecasts on par with a model that is supplied with the true network are achieved, providing a step toward explainable and scalable forecasting of complex systems.

Jonas Braun, Fabian Fischbach, Daniel Köglmayr et al. · 0 citations
Review Open access May 2026

Data-Driven Identification of Stochastic Dynamical Systems

Identifying stochastic dynamical systems from observational data remains a major challenge in applied mathematics and engineering, particularly when complex systems are influenced by random perturbations and incomplete empirical information. This comprehensive review aims to examine state-of-the-art data-driven methods for discovering governing equations, estimating parameters, and predicting the behavior of stochastic dynamical systems. The review systematically analyzes key methodological approaches, including Sparse Identification of Nonlinear Dynamics (SINDy), Dynamic Mode Decomposition (DMD) and its extensions, Koopman operator theory, neural ordinary differential equations, and Bayesian inference. Each approach is evaluated in terms of its theoretical foundations, computational requirements, robustness to noise, and applicability to different classes of stochastic systems. Drawing on numerical experiments and real-world case studies, the findings show that no single method consistently outperforms others across all scenarios. Instead, hybrid approaches that integrate physics-informed constraints with machine learning demonstrate the strongest potential for advancing data-driven system identification. The review concludes that future research should address real-time identification, uncertainty quantification, and the integration of multi-fidelity data sources to improve the reliability and scalability of stochastic system modeling. This work contributes a comprehensive framework for guiding researchers and practitioners in selecting and implementing appropriate identification methods for stochastic dynamical systems.

Rishav Jha, Kameshwar Sahani, S. K. Sahani et al. · 0 citations
Preprint Jul 2026

Single-Snapshot Inference of Network Couplings from Universal Dynamics at Relative Equilibrium

Many real-world systems can be modelled as complex networks whose collective behaviour is governed by hidden interactions between nodes. Existing methods for inferring these interactions typically require controlled perturbations, time-resolved observations or multiple independent snapshots, all of which are often unavailable in practice. Here we show that class-based coupling strengths can be inferred from a single snapshot of node states when the system is observed close to a relative equilibrium. In this regime, all nodes share a common velocity, which can be absorbed into an effective class bias, transforming the inverse problem into a homogeneous linear system. The coefficients of this linear system are determined entirely by the observed local neighbourhoods and their coupling mechanism, enabling the application to arbitrary known coupling functions. We validate the approach on three different linear and nonlinear dynamical systems, recovering relative class-based couplings and, in special cases, absolute couplings. These results show that spatial heterogeneity can substitute for temporal sampling, enabling single-snapshot inference of hidden coupling strengths in networked dynamical systems.

Moritz Lampert, Dominic Grün, Ingo Scholtes · 0 citations
Preprint Aug 2026

Interpretable Causal Discovery via Causal-Effect Constraints

This work considers the task of conditional causal discovery as a Bayesian inference problem, in which the posterior is targeted over causal graphs and parameters conditional on an event such as a causal-effect constraint, and adapts rare-event estimation techniques to perform inference the joint graph-parameter space.

Cixuan Zhang, Guy Van den Broeck, Benjie Wang · 0 citations
Preprint Aug 2026

GENESIS: Towards Explainable Causal Discovery

Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to resolve structural ambiguities in low-sample regimes. Second, although LLM-assisted hybrid approaches improve structure recovery through semantic reasoning, the influence of that reasoning on individual edge decisions remains largely opaque. Consequently, existing hybrid methods fail to satisfy a fundamental requirement: explaining why a particular edge is included or excluded in the learned directed acyclic graph (DAG). This is critical in real-world applications, where no ground-truth DAG exists and every structural decision must be independently justified. We formalize this requirement as decision traceability, requiring every inferred edge to be supported by auditable statistical evidence, Markov Blanket consistency, or explicit domain reasoning. We propose GENESIS, an explainable hybrid CD framework that decomposes graph construction into interpretable decision points. GENESIS first identifies and scores three-node structural motifs, including chains, forks, and colliders, to establish transparent structural priors, then progressively refines the graph by integrating these priors with observational evidence, invoking domain knowledge only when statistical evidence is insufficient. By design, every edge decision is resolved through an auditable source of evidence. Experiments show that GENESIS achieves 100% decision traceability across all settings, establishing explainability as a first-class objective in causal discovery. Despite this additional requirement, GENESIS consistently outperforms purely statistical CD methods on the majority of benchmark datasets across all sample regimes in terms of Structural Hamming Distance (SHD), while achieving performance comparable to state-of-the-art LLM-assisted approaches.

Abhinav Thorat, Ravi Kolla, Vishak K Bhat et al. · 0 citations