STRUMP-I: Structure-Based Machine Learning Approach to pMHC-I Binding Prediction Using Force Field Energy Features
The adaptive immune system monitors cellular integrity by recognizing short peptides from intracellular proteins presented on major histocompatibility complex class I (MHC-I) molecules, collectively termed peptide–MHC complexes (pMHC), enabling detection of foreign or mutated proteins. With the rising importance of immunotherapies targeting cancer neoantigens, accurately predicting which peptides bind to MHC alleles is critical. Current computational methods for pMHC-I binding prediction fall into sequence-based methods, which rely heavily on large training datasets, and structure-based methods that leverage structural modeling and pMHC binding energetics. Although sequence-based methods are widely used, their performance depends on the size and quality of the training data. Structure-based approaches, by contrast, can generalize better across diverse MHC alleles, but they traditionally depend on identifying a single global minimum-energy conformation, an assumption that may be inadequate for the promiscuous binding of MHC-I molecules. To address these limitations, we developed STRUMP-I (STRUcture-based pMHC Prediction for class I), a novel pMHC-I binding prediction tool that directly leverages a broad set of force-field-derived energy terms as machine learning features. In the standard benchmark set, STRUMP-I achieved performance comparable to state-of-the-art sequence-based models overall and showed a clear advantage for alleles with limited or imbalanced representation. Furthermore, STRUMP-I complemented sequence-based methods by removing method-specific false positives and improving precision, with a more favorable precision–recall tradeoff than AF-FT. These evaluations reinforced the value of STRUMP-I as a structure-informed prioritization method, particularly for underrepresented alleles and as a high-precision post-prediction filter.