Skip to content
Open access

Pi-Ensemble: Sequence-guided generation of interpolated protein conformational ensembles

Aug 2026 · bioRxiv · 0 citations
Biology

TL;DR

Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states, provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.

Abstract

Proteins are critical biomolecular machines that populate ensembles of interconverting conformations. Many biological processes depend on transitions between metastable states. Although molecular dynamics (MD) simulations provide a physically grounded route to characterize these motions, routine sampling of large-scale conformational transitions remains computationally demanding. Recent advances in protein structure prediction have created new opportunities for ensemble generation, but many existing approaches require noising inputs, task-specific training, supervised fitting on extensive MD data, or experimentally-informed restraints. Here, we introduce Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states. Unlike previous methods, Pi-Ensemble alternately leverages inverse-folding and structure-prediction models to propose intermediate conformations between known protein states, generating diverse ensembles without additional training. We evaluate Pi-Ensemble across diverse protein systems, including enzymes, transporters, receptors, and benchmark cases with reference MD simulations or experimental Double Electron-Electron Resonance (DEER) data. Pi-Ensemble recovers physically plausible intermediate conformations, captures transition pathways observed in large-scale MD simulations, and generates structures consistent with experimental distance distributions. Furthermore, Pi-Ensemble-generated conformations provide effective starting seeds for parallel MD simulations, improving conformational exploration and accelerating convergence relative to simulations initiated only from endpoint structures. These results establish sequence-guided structural interpolation as a practical strategy for probing protein conformational landscapes. By generating diverse and physically reasonable conformational proposals without long-timescale MD or model retraining, Pi-Ensemble provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.

Read PDF

Similar papers

Open access Aug 2026

Benchmarking AI-generated structural ensembles of membrane proteins against physics-based modelling

It is demonstrated that BioEmu can generate plausible conformational ensembles for relatively large, six-and seven-pass membrane proteins, sampling rare states at a fraction of the computational cost of conventional MD simulations, suggesting that AI-based ensemble generation could provide an accessible approach for exploring membrane protein dynamics and complement conventional molecular modelling approaches.

B. Clifton, Adam G Grieve, Robin A. Corey · 0 citations
Open access Aug 2026

Exploring Conformational Transitions of Adenine RNA Dimer via Machine Learning Potentials

This work assesses ML potentials for exploring RNA conformations using the adenine–adenine dinucleoside monophosphate (ApA) dimer, a fundamental RNA building block, and parametrized ML potentials based on the equivariant MACE architecture and informed by both ab initio and semiempirical property data.

Leonardo Medrano Sandonas, Macarena Tolmos Nehme, L. F. Cofas-Vargas et al. · 0 citations
Open access Jul 2026

Carbonara: a SAXS-guided seeding framework for exploring protein solution-state dynamics

Proteins in solution often populate conformational ensembles that differ from the static states captured by crystallography or AI-based structure prediction. Conventional molecular dynamics (MD) simulations often fail to cross the energy barriers separating these states on accessible timescales, and statistical reweighting cannot recover conformations never sampled. Here we present Carbonara, a framework that uses experimental small-angle X-ray scattering (SAXS) data to predict alternative physically plausible protein conformations. Carbonara builds on Wiggle, a standalone Cα-based SAXS forward model validated against explicit-solvent calculations and experimental benchmarks. Using two case studies, an AI-predicted multi-domain helicase (SMAR-CAL1) and a crystallographic antibody fragment (ChiLob7/4 IgG2), we show how seeding MD simulations from Carbonara conformations enables efficient exploration of solution-state conformational landscapes. In both cases, MD ensembles initiated from available models either fail to match the SAXS data or do so only after discarding nearly all sampled conformations, whereas Carbonara-seeded ensembles reach agreement while retaining the majority of conformations. Our modelling framework provides a route from static structural models of flexible multi-domain proteins and multimeric assemblies to solution-state ensembles.

Josh McKeown, Cameron Brown, Arron Bale et al. · 0 citations
Review Open access Jul 2026

Predicting Biomolecular Interactions in the Next Decade: Physics-Based Methods Meet AI-Driven Approaches.

It is argued that, since physics-based simulations and machine learning provide complementary approximations to the underlying probability distribution associated with biomolecular recognition events, and they excel respectively in consistency with free-energy landscapes and state populations and in predictive accuracy, the central challenge for the coming decade will be integrating them into hybrid frameworks that are scalable and transferable.

R. Khalil, Elena Frasnetti, Han Kurt et al. · 0 citations
Open access Aug 2026

Inferring protein ensembles directly from NOESY spectra

A quantitative scoring framework for comparing experimental and back-calculated observables is introduced and combined with regularized ensemble selection and Monte Carlo simulated annealing to provide direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables.

Murray Coles · 0 citations