Sparse experimental data often yields vast mechanistic hypothesis spaces with numerous equally probable models. Traditional model selection metrics like the Akaike Information Criterion fall short because they reduce models’ nonlinear dynamics to a scalar score that masks crucial mechanistic details. Here, we introduce an AI-driven framework treating model dynamics as learnable signatures. Using deep learning autoencoders, we embed the dynamic signatures of thousands of competing models into a low-dimensional latent space. Iterative clustering and physiological constraints systematically refine this space to a manageable number of testable hypotheses. Applying this to the Integrated Stress Response, a fundamental cellular homeostasis mechanism, we converged on 12 compelling mechanistic hypotheses from over 12,000 candidates. Notably, this framework confidently rejected structure-derived kinetic assumptions regarding higher-order PKR activation that traditional metrics could not resolve. Ultimately, this method provides a rigorous, full-state alternative to scalar metrics, accelerating the discovery of driving principles in complex biological systems.
Mustafa Ozen, Chaitra Agrahar, Francesca Zappa et al.· bioRxiv· 0 citations
Recent benchmarks show that deep-learning models for perturbation prediction do not outperform simple baselines operating in principal-component (PCA) space. We explain this with an information-theoretic ceiling: for any orthonormal projection basis Φ, the squared correlation between prediction and truth is bounded by the variance the basis explains (r2 ≤ VE), so no model complexity can recover signal discarded at projection. On chemical perturbations (sciPlex3, LINCS L1000), the eigenbasis of a gene association network captures only 10–12% of drug-response variance and yields chance-level predictions, while PCA captures 90–99%. Graph wavelets built on the same network recover ∼88%, localising the drug signal in high-frequency modes that the standard low-pass eigenbasis discards. On CRISPRa genetic perturbations the ranking inverts: the network basis outperforms PCA across all dimensions tested. Controls on topology, null networks and data leakage confirm the effect is structural. The right basis depends on the perturbation modality: PCA captures the variance that drives chemical responses, the network basis captures the cascade structure that drives genetic ones, and bases that access the network’s full graph spectrum (such as graph wavelets) recover both from the same topology.