The results show that current LLMs capture partial epitope-related signals but remain limited in antibody-specific sequence grounding, long-context residue localization, and biologically grounded reasoning, so EpiBench provides a diagnostic testbed for measuring and improving sequence-aware biomedical LLMs toward reliable LLM-assisted antibody discovery.
Abstract
Epitopes determine where antibodies bind antigens and shape downstream therapeutic properties such as functional blockade and escape resistance, making epitope understanding central to antibody drug discovery. Although large language models (LLMs) have shown strong biomedical reasoning ability, it remains unclear whether they can infer epitope information directly from antigen and antibody sequences. Existing epitope resources typically focus on isolated prediction tasks or rely on specialized structural settings, while general protein benchmarks do not evaluate epitope-centered decisions across the antibody development workflow. To address this gap, we introduce EpiBench, a closed-book, sequence-based, and automatically scorable benchmark for evaluating epitope reasoning in LLMs. EpiBench contains 1,609 curated samples grounded in structural antibody--antigen contacts, curated functional B-cell assays, and deep mutational scanning escape measurements. It covers five connected tasks: targetable region discovery, antibody-conditioned epitope identification, epitope binning, functional epitope assessment, and antibody escape assessment, with controlled sampling to reduce shortcut-based evaluation artifacts. We evaluate nine general-purpose LLMs and analyze their behavior through task-specific baselines, antigen length stratification, explicit-reasoning comparison, and failure-mode inspection. The results show that current LLMs capture partial epitope-related signals but remain limited in antibody-specific sequence grounding, long-context residue localization, and biologically grounded reasoning. Therefore, EpiBench provides a diagnostic testbed for measuring and improving sequence-aware biomedical LLMs toward reliable LLM-assisted antibody discovery.
SAASBench provides a framework for evaluating the model's ability to estimate the specificity of a candidate antibody in relevant settings, indicating that strong performance on traditional affinity benchmarks does not automatically translate into reliable antibody specificity estimation in proteome-derived settings.
Dmitriy Umerenkov, Ivan Poddiakov· Proceedings of the 32nd ACM...· 0 citations
These findings provide practical guidance for integrating open-source protein structure prediction models into AI-driven nanobody discovery pipelines while highlighting the need for improved generalization across antigens.
Yannick Vogt, Rebekka Roßberg, Jan Habermann et al.· Frontiers in Bioinformatics· 1 citation
Cancer epitopes, the molecular structures recognized by T and B cells at the tumor interface, are central to understanding antitumor immunity and developing immunotherapies. Yet despite the rapid growth of cancer immunology data, a comprehensive, continuously updated, and accessible resource for cancer epitope data has been lacking. The Cancer Epitope Database and Analysis Resource (CEDAR, cedar.iedb.org) was established in 2021 to fill this gap, providing curated experimental epitope data alongside a suite of cancer-specific computational tools for epitope prediction and analysis. Built on the validated infrastructure of the Immune Epitope Database (IEDB), CEDAR integrates cancer epitope data with biological, immunological, and clinical context, enabling researchers to explore immune recognition of tumors, identify candidate targets for immunotherapy, and benchmark prediction methods. Here we describe CEDAR’s current capabilities, report on progress in curation, database development, and tool availability, and outline the opportunities and challenges ahead for expanding its scope and utility to the cancer research community.
Zeynep Koşaloğlu-Yalçın, Ibel Carri, Daniel Marrama et al.· Frontiers in Oncology· 0 citations
ABSTRACT Antibodies are renowned for their ability to bind diverse targets with high affinity and specificity, yet identifying binders with predefined epitope specificity remains a major challenge. In this study, we investigate the concept of mimic antibodies–antibodies that recapitulate the binding mode of a target’s cognate ligand. Through a systematic analysis of the Protein Data Bank (PDB), we show that such mimicry is widespread and arises through diverse structural mechanisms, such as single-loop, multi-loop and scattered interaction mimicry. These findings indicate that protein interfaces impose strong constraints on binding, leading to convergent interaction solutions that can be independently discovered by antibodies. Building on these findings, we developed a ligand-guided strategy to mine immune repertoire data by selecting antibodies whose predicted binding interfaces mimic the interaction motif of a cognate ligand. Applied to the interaction between interleukin-18 (IL-18) and its receptor alpha (IL-18RA), mimicry-guided screening of a 20,000-sequence repertoire yielded 31 candidates, 11 of which (35% hit rate) bound the IL-18RA D3 domain, with eight reaching sub-nanomolar affinities that surpass the cognate ligand. Our findings establish mimic antibodies as a promising strategy for rational antibody selection, engineering, and design, with broad implications for therapeutic antibody development and drug discovery.
Brennan Abanades, J. López-Morales, Ivana Tanasijević et al.· mAbs· 0 citations
The AIntibody challenge shows that AI can optimize antibodies in defined, biologically grounded regimes, in addition to highlighting critical gaps including affinity prediction and library-inspired antibody design and cross-task generalization.
M. Erasmus, Daniel Bedinger, Elizabeth Hopkins et al.· Nature Biotechnology· 0 citations