Skip to content
Preprint

PHASE: encoding global protein ensembles with local Hamiltonians and all-atom backmapping

Aug 2026 · 0 citations · 44 references
Biology

TL;DR

PHASE (Protein Hamiltonians for Sampling of Ensembles), a system-specific framework that converts atomistic conformational ensembles into an explicit and interpretable statistical model, is introduced.

Abstract

Protein function is governed by conformational ensembles, which can be viewed as high-dimensional probability distributions over molecular conformations. Yet the statistical organization of these distributions is often represented only implicitly, either through collections of simulation trajectories or within high-capacity generative models. Here, we introduce PHASE (Protein Hamiltonians for Sampling of Ensembles), a system-specific framework that converts atomistic conformational ensembles into an explicit and interpretable statistical model. Applied to ten conformational ensembles derived from approximately 37$\mu$s of atomistic simulations of the adenosine A2A receptor, Hamiltonians containing only local residue couplings within 6$\mathring{A}$ reproduce residue-wise and pairwise microstate statistics, including correlations between residues that are not directly coupled in the model. Moreover, independently fitted inactive and active reference Hamiltonians define an endpoint preference coordinate that organizes newly sampled ligand-, effector- and conformation-dependent ensembles along the A2A activation landscape without receiving these biochemical labels as model inputs. Finally, a cluster-conditioned all-atom reconstruction model preserves the prescribed residue microstate patterns of newly sampled configurations, closing the coarse-graining-sampling-backmapping cycle. The resulting discrete representation additionally admits direct QUBO encoding, enabling classical annealing and providing a route toward future quantum-annealing implementations. PHASE therefore provides a protein-general procedure for constructing compact, interpretable and atomistically realizable statistical models of protein conformational ensembles.

View source

Similar papers

Open access Aug 2026

Pi-Ensemble: Sequence-guided generation of interpolated protein conformational ensembles

Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states, provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.

Hassan Nadeem, D. Kleiman, Yuming Zhou et al. · 0 citations
Open access Aug 2026

Inferring protein ensembles directly from NOESY spectra

A quantitative scoring framework for comparing experimental and back-calculated observables is introduced and combined with regularized ensemble selection and Monte Carlo simulated annealing to provide direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables.

Murray Coles · 0 citations
Preprint Aug 2026

A Quantum Circuit Framework for Protein Ensemble-Level Energetics

This framework expands quantum protein modelling beyond single-structure optimisation toward ensemble-level characterisation, capturing key features of rugged energy landscapes to guide protein design, mutation mapping, and allosteric pathway identification.

Pratik Patil, Bhushan Bonde, B. Choubey · 0 citations
Jul 2026

Local Geometry Recovers but Cooperative Structure Does Not: Residue-Resolved Limits of All-Atom Reconstruction from Single-Bead Coarse-Grained Disordered Protein Ensembles

Intrinsically disordered proteins (IDPs) drive diverse cellular processes through broad conformational ensembles, but experimental characterization of these ensembles is sparse and high-quality training data is scarce, holding back artificial intelligence (AI) approaches to ensemble prediction. Single-bead coarse-grained (CG) force fields such as CALVADOS, parametrized directly against experimental observables, currently provide a more reliable route to disordered ensembles than direct AI prediction. CG sampling lacks atomistic resolution and must be paired with backmapping; the atomistic information recoverable from this two-step process is shaped jointly by the CG representation and the backmapping algorithm, and the interplay between these contributions is not well characterized. We benchmarked CODLAD, a recently published latent-diffusion backmapping pipeline, on 23 Protein Ensemble Database (PED) systems using a dual-input design: the same architecture receives either PED-reference Cα coordinates or independently sampled CALVADOS Cα trajectories. This design separates paired reconstruction error, measurable for PED+CODLAD, from unpaired CG-input-associated ensemble deviations, measurable for CG+CODLAD at the distribution level. Reconstruction from PED conformers achieved 0.56 ± 0.08 Å backbone root-mean-square deviation and preserved local geometry. Reconstruction from CALVADOS-sampled Cα trajectories preserved the Cα framework and bond geometry, while ensemble-level deviations relative to PED were localized to proline backbone geometry and cooperative secondary structure, both consistent with information not carried by an unconstrained single-bead representation. The benchmark quantifies which atomistic properties can be recovered after projecting a CG ensemble into all-atom space and identifies improved CG geometric encoding and sequence-conditioned AI priors as the directions for further progress.

Jianxiang Huang, Xin Qiao, Ning Liu et al. · 0 citations
Open access Jul 2026

De Novo Design of Protein Switches with Diffusion-Based Ensemble Sampling

Protein switches are proteins that can respond to biochemical stimuli by rearranging their structural elements, essential for cells to transduce signals. The de novo design of such proteins requires amino acid sequences whose energy landscapes support multiple stimulus-dependent conformations, yet most current de novo protein design pipelines are optimized for single stable structures. Existing multi-state inverse-folding methods can design sequences compatible with multiple backbones, but they assume that suitable backbone ensembles are already available, often requiring expert knowledge. We introduce Diff-Switch, a framework for sampling switch-like backbone ensembles from pretrained protein diffusion models. Given a reference backbone structure and domain decomposition, our method preserves local domain geometry while encouraging diversity in global domain arrangements along user-specified collective variables, inspired by metadynamics. We implement this objective through a controlled diffusion sampler with reward-tilting for local similarity between the ensemble members and history-dependent bias in collective-variable space to avoid repeated sampling of the same global arrangement. The resulting ensembles provide candidate conformational states for downstream multi-state inverse folding. Across our evaluation set of 20 diverse proteins, using conformations from these generated ensembles improves the success rate of finding switch-compatible sequences over baseline sampling. We further apply the method to a real-world protein switch design task and characterize the resulting designs.

Alireza Omidi, Jiajun He, Jennifer M. Bui et al. · 0 citations