Skip to content
Open access

Inferring protein ensembles directly from NOESY spectra

Aug 2026 · bioRxiv · 0 citations · 12 references
Biology

TL;DR

A quantitative scoring framework for comparing experimental and back-calculated observables is introduced and combined with regularized ensemble selection and Monte Carlo simulated annealing to provide direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables.

Abstract

Solution NMR spectroscopy provides atomistic measurements of proteins in a native-like biophysical state. Because these measurements are ensemble averages, it also has the potential to report on conformational diversity. However, conventional NMR structure determination typically converts experimental observables into restraints for molecular dynamics, which encode information on the mean structure but do not retain information on the underlying conformational distribution. Ensemble selection has long been proposed as an alternative, whereby experimental observables are compared directly with candidate conformers generated independently of the measurements. This allows population distributions to be inferred from the data. However, few such methods have incorporated NOESY - the richest source of structural information in protein NMR - data, due to challenges in the quantitative comparison of experimental and back-calculated spectra. To address this challenge, we previously introduced the CoMAND method, demonstrating that quantitative agreement is practical for NOESY spectra with bespoke heteronuclear editing schemes. Here we extend this approach into a framework for direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables. We introduce a quantitative scoring framework for comparing experimental and back-calculated observables and combine it with regularized ensemble selection and Monte Carlo simulated annealing. Integration with the OpenMM molecular dynamics engine allows conformational pools to be generated using established molecular simulation methods. Applied to human ubiquitin, the resulting ensemble provides simultaneous agreement with NOESY, residual dipolar coupling and scalar coupling data while retaining conformational diversity supported by experiment.

Read PDF

Similar papers

Open access Aug 2026

Pi-Ensemble: Sequence-guided generation of interpolated protein conformational ensembles

Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states, provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.

Hassan Nadeem, D. Kleiman, Yuming Zhou et al. · 0 citations
Open access Aug 2026

SPINDLE: Unlocking protein dynamics from single-field NMR relaxation data using a deep learning ensemble

A protein’s function is derived from its three-dimensional structure and the motions of the atoms about that structure. The detailed characterization of both macromolecular structure and dynamics provides an opportunity for understanding enzyme catalysis, ligand binding, and allostery, along with providing insights into how the function changes upon mutation or post-translational modification. Among the various methods for characterizing biomolecular motions, nuclear magnetic resonance (NMR) spin relaxation methods are a standard for determining nanosecond global tumbling times along with the amplitude and timescale of faster local motions. Within the model-free formalism, various mathematical models are used to extract dynamic parameters. Unfortunately, as the number of fitted parameters increases within these models, they become mathematically underdetermined for standard NMR relaxation data collected at a single magnetic field, necessitating multi-field datasets. Here, we present SPINDLE, an ensemble of deep neural networks trained on a large synthetic set of NMR relaxation data. Unlike traditional least-squares fitting, SPINDLE predicts both fast and slow timescale dynamics parameters from a single set (i.e., collected at a single magnetic field) of three relaxation datasets using the ensemble for error estimation. We demonstrate a strong correlation to ground truth dynamics parameters on synthetic benchmarks, with more precision than traditional fitting techniques, and precisely reproduce experimental dynamics parameters for ∼50 proteins with relaxation data in the Biological Magnetic Resonance Data Bank. We also leverage the architecture of the deep neural network to show how the model emphasizes rigid residues for the prediction of global correlation times. This strategy may be useful in the future for elucidating correlated networks of dynamic residues from multiple relaxation datasets.

Olivia E. Krise, Michael P. Latham · 0 citations
Open access Jul 2026

Integrative Ensemble Modeling reveals RNA conformations targetable by small molecules

RNA molecules explore heterogeneous conformational ensembles that are essential for their biological function and molecular recognition, yet this intrinsic flexibility poses a major challenge for structure-based drug discovery. In particular, the absence of well-defined binding pockets in static structures limits the identification of ligandable sites. Here, we present an integrative ensemble-based approach that combines enhanced-sampling molecular dynamics simulations with Nuclear Magnetic Resonance data to characterize the conformational landscape of the HIV-1 TAR RNA at atomic resolution. Starting from extensive sampling, we refined the resulting conformational distribution through maximum-entropy reweighting to achieve quantitative agreement with experimental data. Analysis of the reweighted ensemble reveals a diverse set of conformational substates, including compact arrangements that exhibit pocket features compatible with ligand recognition and overlap with known ligand-bound structures. At the same time, highly ligandable conformations, which are only marginally populated, might nonetheless be critical for RNA recognition. Our results demonstrate that integrative ensemble modeling can reveal pharmacologically relevant RNA conformations that are not apparent from experimental static structures, providing a framework for ensemble-based strategies in RNA-targeted drug discovery.

Stefano Bosio, Vincent Schnapka, Mattia Bernetti et al. · 0 citations
Open access Aug 2026

makeshift: a lightweight software for accessing and analyzing NMR data and protein dynamics

Nuclear magnetic resonance (NMR) spectroscopy yields rich residue-level information on biomolecular dynamics and chemical environments, two frontiers for quantitative predictive methods in biochemistry. Decades of data are publicly archived in the Biological Magnetic Resonance Data Bank (BMRB)1, yet in practice, this information remains difficult to access and interpret at scale and within computational workflows. Here we present makeshift, an open-source Python package for accessing, curating, and analyzing NMR datasets. Users can readily retrieve and parse BMRB entries and perform essential analyses such as chemical shift re-referencing, secondary structure propensity prediction, and interpretation of relaxation datasets for dynamics. We re-implemented several widely-used NMR data calculations which were not open-source or available in Python and validated our implementations against the original implementations. By integrating data access, processing, and analysis into a single Python interface, makeshift lowers the barrier for reproducible, scalable analysis and machine learning applications using biomolecular NMR data.

Gina El Nesr, Hannah K. Wayment-Steele · 0 citations
Open access Jul 2026

Carbonara: a SAXS-guided seeding framework for exploring protein solution-state dynamics

Proteins in solution often populate conformational ensembles that differ from the static states captured by crystallography or AI-based structure prediction. Conventional molecular dynamics (MD) simulations often fail to cross the energy barriers separating these states on accessible timescales, and statistical reweighting cannot recover conformations never sampled. Here we present Carbonara, a framework that uses experimental small-angle X-ray scattering (SAXS) data to predict alternative physically plausible protein conformations. Carbonara builds on Wiggle, a standalone Cα-based SAXS forward model validated against explicit-solvent calculations and experimental benchmarks. Using two case studies, an AI-predicted multi-domain helicase (SMAR-CAL1) and a crystallographic antibody fragment (ChiLob7/4 IgG2), we show how seeding MD simulations from Carbonara conformations enables efficient exploration of solution-state conformational landscapes. In both cases, MD ensembles initiated from available models either fail to match the SAXS data or do so only after discarding nearly all sampled conformations, whereas Carbonara-seeded ensembles reach agreement while retaining the majority of conformations. Our modelling framework provides a route from static structural models of flexible multi-domain proteins and multimeric assemblies to solution-state ensembles.

Josh McKeown, Cameron Brown, Arron Bale et al. · 0 citations