Scoring biomolecular complexes is central to structure assessment and drug discovery, yet the complexes themselves vary widely in pose, size, and molecular composition. A scoring function tuned for one interaction type rarely carries over to another, and most existing methods compound the problem by leaning heavily on task-specific labels. We introduce OmniScore, a universal structure-based framework that learns a shared geometry-aware representation of complexes once and then adapts it to downstream scoring through lightweight task-specific heads. OmniScore couples a graph view and a sequence view of each structure, encodes its three-dimensional geometry, and compresses representations into a compact latent space that a reconstruction module and prediction heads can reuse. We pretrain this backbone on diverse datasets including complexes, monomers, and small molecules with complementary objectives: coordinate recovery, correcting corrupted input tokens, predicting molecular identity, and grounding the representation in structure-level physical quantities. Across the evaluated benchmarks, OmniScore gave the best antibody-antigen and nanobody-antigen quality assessment on all reported metrics compared to state-of-the-art baselines. Its frozen residue embeddings matched the state-of-the-art protein-tokenization method with an average functional-site accuracy of 71.8% on a standard residue-level benchmark. On protein-ligand scoring and ranking benchmarks, it performed on par with methods built specifically for that single task. These results suggest that geometry-aware pretraining can provide a reusable scoring backbone for tasks that depend on interfacial and residue-level structure, within the evaluated settings.
The adaptive immune system monitors cellular integrity by recognizing short peptides from intracellular proteins presented on major histocompatibility complex class I (MHC-I) molecules, collectively termed peptide–MHC complexes (pMHC), enabling detection of foreign or mutated proteins. With the rising importance of immunotherapies targeting cancer neoantigens, accurately predicting which peptides bind to MHC alleles is critical. Current computational methods for pMHC-I binding prediction fall into sequence-based methods, which rely heavily on large training datasets, and structure-based methods that leverage structural modeling and pMHC binding energetics. Although sequence-based methods are widely used, their performance depends on the size and quality of the training data. Structure-based approaches, by contrast, can generalize better across diverse MHC alleles, but they traditionally depend on identifying a single global minimum-energy conformation, an assumption that may be inadequate for the promiscuous binding of MHC-I molecules. To address these limitations, we developed STRUMP-I (STRUcture-based pMHC Prediction for class I), a novel pMHC-I binding prediction tool that directly leverages a broad set of force-field-derived energy terms as machine learning features. In the standard benchmark set, STRUMP-I achieved performance comparable to state-of-the-art sequence-based models overall and showed a clear advantage for alleles with limited or imbalanced representation. Furthermore, STRUMP-I complemented sequence-based methods by removing method-specific false positives and improving precision, with a more favorable precision–recall tradeoff than AF-FT. These evaluations reinforced the value of STRUMP-I as a structure-informed prioritization method, particularly for underrepresented alleles and as a high-precision post-prediction filter.
Adam Voshall, Jeongjun Chae, Honglan Li et al.· Computational and Structural...· 0 citations