It is demonstrated that encoding conformational ensembles into a single thermodynamically informed embedding improves cyclic-peptide property prediction.
Abstract
Molecular property prediction from structure often uses a single representative conformation, even though many molecules exist as conformational ensembles in solution. We introduce EnsembleEGNN, a molecular ensemble foundation model that encodes an ensemble by first encoding each conformer with shared Equivariant Graph Neural Network (EGNN) layers, then pooling the resulting conformer representations with a Set Attention Block. We pretrain the model on CREMP, a cyclic peptide ensemble dataset, using a multi-task self-supervised objective combining masked token recovery, noisy-coordinate reconstruction, and pairwise distance reconstruction. On the CREMP-CycPeptMPDB dataset, training EnsembleEGNN from scratch fails entirely ($R^2=0.005$). However, the pretrained model reaches $R^2=0.477$ and Pearson $r=0.699$, outperforming the sequence-only BERT baseline ($R^2=0.439$, Pearson $r=0.667$). When EnsembleEGNN is co-trained end-to-end with the BERT sequence encoder, the hybrid model improves further to $R^2=0.538$ and Pearson $r=0.737$. These results demonstrate that encoding conformational ensembles into a single thermodynamically informed embedding improves cyclic-peptide property prediction.
PHASE (Protein Hamiltonians for Sampling of Ensembles), a system-specific framework that converts atomistic conformational ensembles into an explicit and interpretable statistical model, is introduced.
Daniele Angioletti, Marco S. Nobile, Matteo Carli et al.· 0 citations
The SBMR-CNN model demonstrates highly competitive accuracy, outperforming the CM, Uni-Mol+, and MPNN-2D benchmarks, while closely approaching the performance of the more computationally intensive MPNN-3D and SOAP descriptors, as well as the RF-MF model.
Abdulaziz W. Alherz, C. Tezak, Mohammed S. Alhajeri· Industrial & Engineering...· 0 citations
Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states, provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.
Hassan Nadeem, D. Kleiman, Yuming Zhou et al.· bioRxiv· 0 citations
MultiGeo is a DTA prediction framework that explicitly leverages multiple protein conformations rather than a single snapshot, and introduces a disagreement-aware gating mechanism that adaptively fuses this ensemble representation with the dominant structure only when the additional conformers provide complementary information.
Rui-Da Zeng, Cheng Guo, Yajie Meng et al.· 0 citations
Scoring biomolecular complexes is central to structure assessment and drug discovery, yet the complexes themselves vary widely in pose, size, and molecular composition. A scoring function tuned for one interaction type rarely carries over to another, and most existing methods compound the problem by leaning heavily on task-specific labels. We introduce OmniScore, a universal structure-based framework that learns a shared geometry-aware representation of complexes once and then adapts it to downstream scoring through lightweight task-specific heads. OmniScore couples a graph view and a sequence view of each structure, encodes its three-dimensional geometry, and compresses representations into a compact latent space that a reconstruction module and prediction heads can reuse. We pretrain this backbone on diverse datasets including complexes, monomers, and small molecules with complementary objectives: coordinate recovery, correcting corrupted input tokens, predicting molecular identity, and grounding the representation in structure-level physical quantities. Across the evaluated benchmarks, OmniScore gave the best antibody-antigen and nanobody-antigen quality assessment on all reported metrics compared to state-of-the-art baselines. Its frozen residue embeddings matched the state-of-the-art protein-tokenization method with an average functional-site accuracy of 71.8% on a standard residue-level benchmark. On protein-ligand scoring and ranking benchmarks, it performed on par with methods built specifically for that single task. These results suggest that geometry-aware pretraining can provide a reusable scoring backbone for tasks that depend on interfacial and residue-level structure, within the evaluated settings.
Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.
Vilya Research Pascal Sturmfels, Naozumi Hiranuma, M. Salem et al.· 0 citations