Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states, provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.
Abstract
Proteins are critical biomolecular machines that populate ensembles of interconverting conformations. Many biological processes depend on transitions between metastable states. Although molecular dynamics (MD) simulations provide a physically grounded route to characterize these motions, routine sampling of large-scale conformational transitions remains computationally demanding. Recent advances in protein structure prediction have created new opportunities for ensemble generation, but many existing approaches require noising inputs, task-specific training, supervised fitting on extensive MD data, or experimentally-informed restraints. Here, we introduce Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states. Unlike previous methods, Pi-Ensemble alternately leverages inverse-folding and structure-prediction models to propose intermediate conformations between known protein states, generating diverse ensembles without additional training. We evaluate Pi-Ensemble across diverse protein systems, including enzymes, transporters, receptors, and benchmark cases with reference MD simulations or experimental Double Electron-Electron Resonance (DEER) data. Pi-Ensemble recovers physically plausible intermediate conformations, captures transition pathways observed in large-scale MD simulations, and generates structures consistent with experimental distance distributions. Furthermore, Pi-Ensemble-generated conformations provide effective starting seeds for parallel MD simulations, improving conformational exploration and accelerating convergence relative to simulations initiated only from endpoint structures. These results establish sequence-guided structural interpolation as a practical strategy for probing protein conformational landscapes. By generating diverse and physically reasonable conformational proposals without long-timescale MD or model retraining, Pi-Ensemble provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.
It is demonstrated that BioEmu can generate plausible conformational ensembles for relatively large, six-and seven-pass membrane proteins, sampling rare states at a fraction of the computational cost of conventional MD simulations, suggesting that AI-based ensemble generation could provide an accessible approach for exploring membrane protein dynamics and complement conventional molecular modelling approaches.
B. Clifton, Adam G Grieve, Robin A. Corey· bioRxiv· 0 citations
This work assesses ML potentials for exploring RNA conformations using the adenine–adenine dinucleoside monophosphate (ApA) dimer, a fundamental RNA building block, and parametrized ML potentials based on the equivariant MACE architecture and informed by both ab initio and semiempirical property data.
Leonardo Medrano Sandonas, Macarena Tolmos Nehme, L. F. Cofas-Vargas et al.· Journal of Chemical Theory a...· 0 citations
PHASE (Protein Hamiltonians for Sampling of Ensembles), a system-specific framework that converts atomistic conformational ensembles into an explicit and interpretable statistical model, is introduced.
Daniele Angioletti, Marco S. Nobile, Matteo Carli et al.· 0 citations
Proteins in solution often populate conformational ensembles that differ from the static states captured by crystallography or AI-based structure prediction. Conventional molecular dynamics (MD) simulations often fail to cross the energy barriers separating these states on accessible timescales, and statistical reweighting cannot recover conformations never sampled. Here we present Carbonara, a framework that uses experimental small-angle X-ray scattering (SAXS) data to predict alternative physically plausible protein conformations. Carbonara builds on Wiggle, a standalone Cα-based SAXS forward model validated against explicit-solvent calculations and experimental benchmarks. Using two case studies, an AI-predicted multi-domain helicase (SMAR-CAL1) and a crystallographic antibody fragment (ChiLob7/4 IgG2), we show how seeding MD simulations from Carbonara conformations enables efficient exploration of solution-state conformational landscapes. In both cases, MD ensembles initiated from available models either fail to match the SAXS data or do so only after discarding nearly all sampled conformations, whereas Carbonara-seeded ensembles reach agreement while retaining the majority of conformations. Our modelling framework provides a route from static structural models of flexible multi-domain proteins and multimeric assemblies to solution-state ensembles.
It is argued that, since physics-based simulations and machine learning provide complementary approximations to the underlying probability distribution associated with biomolecular recognition events, and they excel respectively in consistency with free-energy landscapes and state populations and in predictive accuracy, the central challenge for the coming decade will be integrating them into hybrid frameworks that are scalable and transferable.
R. Khalil, Elena Frasnetti, Han Kurt et al.· Journal of Physical Chemistr...· 0 citations
A quantitative scoring framework for comparing experimental and back-calculated observables is introduced and combined with regularized ensemble selection and Monte Carlo simulated annealing to provide direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables.