It is shown that increased sampling and AlphaFold3 generally improve performance relative to default sampling and AlphaFold2, however predictive accuracy and improvement levels varied considerably among interface classes, with antibody-peptide complexes representing a challenge despite their small antigen size.
Abstract
Determining the structural basis of antigen recognition by antibodies and T cell receptors (TCRs) provides critical insights into effective immune targeting and can inform design of biotherapeutics and vaccines. Accurate computational modeling of antibodies and TCRs in complex with their targets poses a major challenge for predictive methods, including AlphaFold, which is generally accurate for modeling protein complexes but has shown limited success for immune recognition. In this study we assessed the performance of AlphaFold2, AlphaFold3, increased sampling protocols, and related deep learning methods for modeling antibody-protein, antibody-peptide, and TCR-peptide-major histocompatibility complex (pMHC) recognition. We show that increased sampling and AlphaFold3 generally improve performance relative to default sampling and AlphaFold2, however predictive accuracy and improvement levels varied considerably among interface classes, with antibody-peptide complexes representing a challenge despite their small antigen size. Comparing per-case success across methods showed some complementarity, indicating opportunities for increased success through model pooling approaches, for instance increasing antibody-peptide near-native success from 41% to 59%. Analysis of AlphaFold confidence scores and modeling of a noncanonical complex provided further insights into predictive performance. These results highlight considerations for predictive antibody and TCR complex modeling efforts, while revealing key distinctions among protocols, scoring, and immune complex classes.
This review compiles the data resources related to polyreactive antibodies and places a particular emphasis on computational models for predicting antibody polyreactivity, which includes empirical models based on physicochemical properties, traditional machine learning models, deep learning networks, and protein language models.
Haoxian Tang, Zixuan Zhang, Wenzhi Li et al.· Computational Biomedicine· 0 citations
The ImmunoFoundation Model (IFM), a multimodal deep learning system that integrates not only peptide sequences, 3D molecular structures, and biochemical properties but also TCR-MHC-peptide to achieve superior immunogenicity prediction and enable peptide optimization for therapeutic applications is developed.
Smita Krishnaswamy, J. Rocha, Hiren Madhu et al.· Journal of Immunology· 0 citations
Adaptive immunity relies on T-cell receptor (TCR) recognition of non-self epitopes, short peptides presented by the Major Histocompatibility Complex (MHC) on the cell surface. Accurate computational prediction of TCR-epitope binding would unlock the development of targeted immunotherapies, such as cancer vaccines and TCR T cell therapies, while simultaneously deepening our fundamental understanding of self/nonself discrimination, pathogen recognition, and autoimmunity. We created an ensemble approach (enFoldX) that leverages structure prediction models such as AlphaFold3 to build sensitive binding predictors. enFoldX can distinguish T cell reactivity between peptides that differ by a single amino acid substitution, as needed for cancer neoantigens. enFoldX utilizes a customized highly parallelized workflow which allows us to produce ensembles of predicted protein structures at scale and train classifiers to infer reactivity based on distributions of engineered structure features and alignment confidence metrics. While state-of-the-art sequence-based approaches we evaluated could predict well for observed TCRs and epitopes close in sequence to training data, their applicability to novel sequences was limited. Conversely, our ensemble approach is the only model that showed true generalizability to novel datasets and even across species. Moreover, our ensemble approach outperforms the current co-folding methods which rely on predictions from the single top ranked structure. By leveraging the entire protein universe at scale, structure ensembles therefore enable classifiers that reflect physical free energies, providing a tractable path towards TCR T therapy design at the sensitivity required for cancer neoantigen discrimination and imparting lessons for a wide array of complex binding problems.
Olga Lyudovyk, Jonathan A. Levine, Melissa Pathil, Stephen Martis, Yuval Elhanati, Vinod P. Balachandran, Quaid Morris, Benjamin D. Greenbaum. enFoldX: AI classification of AlphaFold3-derived structural ensembles enables T cell specificity prediction [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr A028.
O. Lyudovyk, Jonathan A. Levine, M. Pathil et al.· Clinical Cancer Research· 0 citations
AAMFM, an Antigen-specific Antibody Multimodal Foundation Model that learns unified representations of antibody sequences and structures conditioned on antigen context, achieves state-of-the-art performance in functional antibody design, revealing its potential for antigen-specific antibody engineering.
Xiaoliang Shi, Zichen Wang, Runze Ma et al.· 0 citations
CLDN18.2 is a promising tumor-specific antigen; however, the development of therapeutic antibodies against it is challenged by the need for simultaneous optimization of affinity and developability. To address this, we present cdrGPT, a deep learning framework based on GPT-2 for de novo generation of complementarity-determining region H3 (CDRH3) sequences. Our approach integrates pre-training on the Observed Antibody Space (OAS) database with structural templating derived from the known antibody zolbetuximab. Generated sequences were iteratively refined through rejection sampling and fine-tuned against a multi-parameter objective function encompassing predicted affinity and MHC class II binding risk. From an initial set of 50,000 sequences, this screening pipeline yielded 313 high-confidence candidates. Subsequent analysis using evolutionary scale modeling 2 (ESM2) embeddings, principal component analysis (PCA), and clustering revealed three structurally distinct clusters, with intra-cluster cosine similarities exceeding 0.99. Validation of seven representative sequences from the dominant cluster using AlphaFold3 confirmed high structural fidelity to the zolbetuximab template, demonstrating a root mean square deviation (RMSD) of 1.331 Å for the CDRH3 loop and positional deviations of less than 0.4 Å for key paratope residues. These results indicate that the designed variants preserve the core binding mode of the parent antibody. This study establishes a feasible pipeline for integrating AI-generated CDRH3 loops into functional antibody scaffolds, providing a foundation for the accelerated development of therapeutics targeting CLDN18.2 and other clinically relevant antigens.
Tao Qu, Lingyan Yuan, Weiran Cui et al.· PLoS Computational Biology· 0 citations
Adaptive immunity relies on T-cell receptor (TCR) recognition of non-self epitopes, short peptides presented by the Major Histocompatibility Complex (MHC) on the cell surface. Accurate computational prediction of TCR-epitope binding would unlock the development of targeted immunotherapies, such as cancer vaccines and TCR T cell therapies, while simultaneously deepening our fundamental understanding of self/nonself discrimination, pathogen recognition, and autoimmunity. We created an ensemble approach (enFoldX) that leverages structure prediction models such as AlphaFold3 to build sensitive binding predictors. enFoldX can distinguish T cell reactivity between peptides that differ by a single amino acid substitution, as needed for cancer neoantigens. enFoldX utilizes a customized highly parallelized workflow which allows us to produce ensembles of predicted protein structures at scale and train classifiers to infer reactivity based on distributions of engineered structure features and alignment confidence metrics. While state-of-the-art sequence-based approaches we evaluated could predict well for observed TCRs and epitopes close in sequence to training data, their applicability to novel sequences was limited. Conversely, our ensemble approach is the only model that showed true generalizability to novel datasets and even across species. Moreover, our ensemble approach outperforms the current co-folding methods which rely on predictions from the single top ranked structure. By leveraging the entire protein universe at scale, structure ensembles therefore enable classifiers that reflect physical free energies, providing a tractable path towards TCR T therapy design at the sensitivity required for cancer neoantigen discrimination and imparting lessons for a wide array of complex binding problems.
Olga Lyudovyk, Jonathan A. Levine, Melissa Pathil, Stephen Martis, Yuval Elhanati, Vinod P. Balachandran, Quaid Morris, Benjamin D. Greenbaum. enFoldX: AI classification of AlphaFold3-derived structural ensembles enables T cell specificity prediction [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr PR004.
O. Lyudovyk, Jonathan A. Levine, M. Pathil et al.· Clinical Cancer Research· 0 citations