It is shown that conventional random train-test splits inflate retrieval accuracy by 15–28 percentage points owing to clonal leakage, and clone-aware benchmarking as a practical standard of comparison is established, defining the strengths and limitations of frozen zero-shot PLM embeddings for immune receptor retrieval and establishing clone-aware benchmarking as a practical standard of comparison.
Somatic hypermutation (SHM) involves activation-induced cytidine deaminase (AID)-mediated DNA targeting, followed by error-prone repair which introduces mutations. However, existing computational models do not reflect this, limiting their ability to model the mechanism of affinity maturation in silico. Here, we develop a biologically grounded, transformer-based framework that mirrors SHM. Our framework employs two independent antibody language models (AbLMs): an AID-like targeting model to select mutation sites and a DNA repair-like substitution model to predict resulting amino acids.
We sequenced memory B cell repertoires from eight healthy donors using high-accuracy bulk NGS of unpaired VH/VL chains. These data were used for pretraining the two AbLMs to perform in silico SHM. Using AntiBERTa2-derived metrics as a correlate for native-like antibodies, we compared model-generated sequences to the natural human immune repertoire.
Bulk NGS produced ∼84 million high-quality, productive memory B cell receptor sequences for AbLM training. Model predicted mutations matched the spatial distribution and overall load of true SHM. AntiBERTa2 embedding projections and log likelihood scoring revealed that model-generated sequences closely resemble native antibodies and strongly mimic SHM mutational patterns. Model-generated sequences also exhibited significantly increased levels of expression in HEK293F cells, indicating learning of beneficial mutations that enhance protein fitness.
Our framework models SHM to generate native-like human antibodies, providing a biologically grounded tool for guiding design within the natural sequence space. Unlike motif-based or single-stage SHM simulators, our framework explicitly decouples targeting and substitution to enable faithful reproduction of both hotspot localization and amino acid substitution. This approach advances in silico modeling of humoral immunity and may reveal new mechanistic insights into antibody clonal dynamics.
Endowed Fellowship in the Skaggs Graduate School of Chemical and Biological Sciences, Achievement Rewards for College Scientists (ARCS) Foundation San Diego Chapter, National Institute of Allergy and Infectious Diseases (NIAID)
Computational and Systems Immunology (COMP)
Karenna Ng, B. Némoz, Bryan S. Briney· Journal of Immunology· 0 citations
The B-cell receptor (BCR) repertoire serves as a historical record of immunological events. However, deciphering antigen-specific sequences from this vast dataset remains a challenge, particularly for novel pathogens where prior knowledge is absent. While time-course analysis methods such as QASAS have proven effective for tracking immune responses, they rely on existing antibody databases, limiting their applicability to emerging diseases. To overcome this limitation, we introduce LM-QASAS, a reference-free computational framework that integrates antibody language models with repertoire dynamics. By mapping sequences into a high-dimensional semantic embedding space, LM-QASAS identifies functionally convergent clusters of sequences that are semantically similar and exhibit transient expansion upon immune stimulation. In healthy individuals vaccinated with SARS-CoV-2 mRNA vaccines, our method identified spike-specific sequences with over 90% purity, significantly outperforming methods based on simple sequence identity or abundance. Leave-one-out cross-validation demonstrated that LM-QASAS could accurately reconstruct immune dynamics in unseen individuals without external references. Conversely, the method showed limited sensitivity in an influenza vaccine cohort, revealing that the approach is most effective under conditions of robust clonal expansion (high signal-to-noise ratio), such as those induced by mRNA vaccines. LM-QASAS provides a rapid, high-precision platform for monitoring humoral immunity against emerging threats.
Genki Masuda, Y. Funakoshi, S. Iizumi et al.· medRxiv· 0 citations
Introduction How much of the variable (V) and joining (J) gene identity of a T-cell receptor is recoverable from its third complementarity-determining region (CDR3) amino-acid sequence alone? Immune repertoire studies often report the CDR3 with V and J annotation that is missing, low-confidence, or inconsistent, so what the CDR3 alone can and cannot fix is both a basic question about the receptor and a practical one for reading those repertoires. Methods For each of 118,096 pooled human rearrangements (37,687 α and 80,409 β) we computed the posterior distribution over candidate genes under a generative model of V(D)J recombination and under its post-selection counterpart, and measured recoverability by conditional entropy, the candidate-list size needed to contain the annotated gene, the fraction of sequences admitting a high-confidence single-gene call, and the structure of gene-by-gene confusion. Results The J gene was nearly determined by the CDR3 in both chains. The V gene was only partially recoverable, and behaved as a group rather than a gene: junctional trimming and non-templated insertion, together with the loss of synonymous codon information in translation, leave sets of mutually confusable V genes whose grouping departs sharply from germline family nomenclature (adjusted Rand index 0.05 for α and 0.21 for β). Selection sharpened the V posterior modestly (usage-controlled entropy shift −0.06 nats for α and −0.28 for β) and redistributed which V gene was most probable, a locus-scale rewrite in β against a mild reweight in α. Both the recoverability measurements and the confusion grouping reproduced in two held-out tumor cohorts. Discussion V identity is an emergent, system-level property of the repertoire, set jointly by recombination and selection and invisible in any single rearrangement, so it should be reported as a calibrated group rather than a single gene. We also release the pipeline with a computational tool which can output a set of candidate genes with confidence values given a CDR3 sequence.
The Immune Epitope Database (IEDB, iedb.org) is a freely available resource that catalogs experimentally defined immune epitopes. Concurrently, the IEDB records ∼190,000 T cell receptors and ∼5,000 antibodies with experimentally verified epitope specificity. Because these receptors have been manually curated from 3,300 references spanning decades, reported data and nomenclature can be inconsistent, posing challenges for computational analyses. To support interoperability and integration with community resources such as the Adaptive Immune Receptor Repertoire Knowledge Commons (AKC), we are revising all immune receptor records to produce resolved, standardized, and analysis-ready receptor data.
We developed a computational pipeline that employs IgBLAST for V/D/J gene assignment, ANARCII for identification of Complementarity Determining Regions (CDRs), and tidytcells to standardize author-reported gene names. We furthermore extended tidytcells to validate and standardize CDR3 sequences based on reported V/J gene usage and to support antibody data. Crucially, the pipeline also flags anomalous data for targeted re-curation by expert curators.
The reprocessed receptor dataset contains V/D/J gene names that are correctly formatted and mapped to existing reference genes, and CDR3 sequences are consistently represented up to their conserved anchor residues. Improved anomaly detection allowed us to identify and correct anomalous receptor records from hundreds of studies.
These revisions increase data quality and improve interoperability, as exemplified by integration with the AKC. This integration will enable researchers to seamlessly query large-scale repertoires for receptors with experimentally verified specificity in the IEDB, link orphan sequences to known targets, and support cross-repository studies of receptor-epitope pairs and their relationship to health and disease.
The IEDB is funded by NIAID contract 75N93019C00001.
The AIRR Knowledge Commons is supported by a U24 (U24I177622) from the NIAID.
Computational and Systems Immunology (COMP)
Lonneke Scheffer, Eve Richardson, R. Vita et al.· Journal of Immunology· 0 citations
Adaptive immunity relies on T-cell receptor (TCR) recognition of non-self epitopes, short peptides presented by the Major Histocompatibility Complex (MHC) on the cell surface. Accurate computational prediction of TCR-epitope binding would unlock the development of targeted immunotherapies, such as cancer vaccines and TCR T cell therapies, while simultaneously deepening our fundamental understanding of self/nonself discrimination, pathogen recognition, and autoimmunity. We created an ensemble approach (enFoldX) that leverages structure prediction models such as AlphaFold3 to build sensitive binding predictors. enFoldX can distinguish T cell reactivity between peptides that differ by a single amino acid substitution, as needed for cancer neoantigens. enFoldX utilizes a customized highly parallelized workflow which allows us to produce ensembles of predicted protein structures at scale and train classifiers to infer reactivity based on distributions of engineered structure features and alignment confidence metrics. While state-of-the-art sequence-based approaches we evaluated could predict well for observed TCRs and epitopes close in sequence to training data, their applicability to novel sequences was limited. Conversely, our ensemble approach is the only model that showed true generalizability to novel datasets and even across species. Moreover, our ensemble approach outperforms the current co-folding methods which rely on predictions from the single top ranked structure. By leveraging the entire protein universe at scale, structure ensembles therefore enable classifiers that reflect physical free energies, providing a tractable path towards TCR T therapy design at the sensitivity required for cancer neoantigen discrimination and imparting lessons for a wide array of complex binding problems.
Olga Lyudovyk, Jonathan A. Levine, Melissa Pathil, Stephen Martis, Yuval Elhanati, Vinod P. Balachandran, Quaid Morris, Benjamin D. Greenbaum. enFoldX: AI classification of AlphaFold3-derived structural ensembles enables T cell specificity prediction [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr PR004.
O. Lyudovyk, Jonathan A. Levine, M. Pathil et al.· Clinical Cancer Research· 0 citations