Skip to content
Open access

What Do Generative Models Learn About Adaptive Immune Receptor Repertoires? A Benchmark Study

Jul 2026 · bioRxiv · 0 citations · 60 references
Medicine Biology

TL;DR

A suite of evaluation metrics tailored to AIRR sequence data are applied and a systematic comparison of popular generative model families proposed for the AIRR field, including variational autoencoders, long short-term memory networks, antibody language models, selection models, and simple statistical baselines are presented.

Abstract

Generative models are increasingly used to model adaptive immune receptor repertoire (AIRR) sequence distributions, promising to decode the sequence diversity shaping immune responses and accelerate the design of therapeutic antibodies and T-cell receptors. Yet it remains unclear whether these models produce biologically meaningful outputs or merely capture surface-level sequence statistics while missing features driven by receptor generation and selection. Rigorous evaluation is needed, but the field lacks established standards, as existing machine learning metrics do not all translate directly to the AIRR domain, given the complex structure of the data and the lack of biological ground truth. Consequently, researchers face difficulties in evaluating the models and selecting appropriate ones, which can critically affect downstream clinical applications. Here, we apply a suite of evaluation metrics tailored to AIRR sequence data and present a systematic comparison of popular generative model families proposed for the AIRR field, including variational autoencoders, long short-term memory networks, antibody language models, selection models, and simple statistical baselines. We focus specifically on the task of learning individual-specific immune receptor repertoires, a clinically relevant challenge with direct implications for personalized immunotherapy, disease monitoring, and vaccine response studies. By analyzing the sequences generated by each model, we identify memorization risks, innovation capabilities, and sensitivity to hyperparameter tuning. Taken together, these results advance the understanding of how current generative models reproduce the biology of individual immune repertoires and lay the groundwork for more principled model development and evaluation.

Read PDF

Similar papers

Sep 2026

Deciphering T-cell receptor-antigen recognition through interpretable residue-level interaction modeling.

Accurate identification of interactions between T-cell receptors (TCRs) and antigenic peptides presented by major histocompatibility complex (MHC) molecules is essential for advancing precision immunotherapy. However, existing approaches often exhibit limited generalization to unseen peptides and struggle to capture th...

Wen-Yu Xi, Ruheng Wang, Xiu-Cai Ye et al. · 0 citations
#machine learning Preprint Sep 2026

An immune world model for multiscale forecasting and therapeutic hypothesis generation

Immune therapies act across cell-intrinsic programs, tissue ecosystems, and patient-specific immune states, yet most predictors address these scales separately. We used a governed evolutionary AI Scientist to construct the Immune World Model, an action-conditioned model that learns how interventions move immune states...

Tao-Yong Cui, Xi Wang, Zong-Hang Li et al. · 0 citations
Open access Sep 2026

OmniTCR: a foundation model unifying T cell receptor recognition prediction and conditional sequence generation

T cell receptor (TCR) recognition prediction and receptor generation are traditionally modelled separately, leaving vast TCR sequence collections disconnected from smaller TCR–peptide–MHC datasets. Here we present OmniTCR, a 113-million-parameter autoregressive foundation model pretrained on 328 million formatted human...

Fei-Ran Zeng, Duanyu Feng, Dan-Dan Song et al. · 0 citations
Open access Sep 2026

Generative Language Modeling for Antibody CDR Grafting and Alignment-driven De Novo Design

GenCDR, a family of LLaMa-based autoregressive language models that read all frameworks as a conditioning prompt and generate all CDRs jointly as a variable-length response, makes CDR likelihoods a clean, separable target for reward attribution, achieves the highest CDR recovery among autoregressive models.

Ferran Gonzalez Hernandez, Oliver M. Turnbull, Maliha Sultana et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Explainability from Training with Applications to TCR-Epitope Prediction

Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models organize evidence and evolve during learning. We introduce e...

Jiarui Li, Zi-Xiang Yin, S. Landry et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.