Skip to content
Open access

Folding scFv–Antigen Complexes at Scale

Jul 2026 · bioRxiv · 0 citations · 28 references
Biology

TL;DR

A scalable bench-marking pipeline is introduced that generates large ensembles of scFv–Ag structure predictions by cofolding a curated subset of 3,800 Ab–Ag complexes from SAbDab using multiple state-of-the-art models under diverse inference-time settings and evaluating model performance in recovering correct scFv–Ag interfaces.

Abstract

Accurate modeling of antibody–antigen (Ab–Ag) complexes is central to biologic development, yet the reliability and failures of modern Ab–Ag folding pipelines remain poorly characterized. Single-chain variable fragments (scFvs) are thera-peutically important antibodies, but large-scale evaluations of structure prediction models on scFv–Ag complexes are largely lacking. We introduce a scalable bench-marking pipeline that generates large ensembles of scFv–Ag structure predictions by cofolding a curated subset of 3,800 Ab–Ag complexes from SAbDab using multiple state-of-the-art models under diverse inference-time settings. The resulting dataset, SCALE (scFv–Ag CompLex Ensembles), includes standardized scFv–Ag sequences and around 200,000 predicted complexes spanning different models, sampling strategies, and auxiliary inputs. Using SCALE, we evaluate model performance in recovering correct scFv–Ag interfaces and assess the ability of existing confidence metrics to select the best structure from prediction ensembles. We find that while confidence scores effectively distinguish easy from hard scFv–Ag complexes, they often fail to identify the highest-quality interface for a given target. Further analysis shows that near-correct interfaces typically appear in ensembles but at low frequency, and inference-time choices such as sampling, recycling, and using evolutionary or structural information are crucial for accurate scFv–Ag complex predictions. Dataset and analysis code are available at https://huggingface.co/datasets/ravishah1/SCALE

Read PDF

Similar papers

Open access Aug 2026

Benchmarking antibody-antigen co-folding on human monomeric antigens

The most recent methods substantially outperformed earlier ones, producing medium-or-better top-ranked models for approximately half of post-cutoff Fv complexes without templates or experimental restraints, and performing similarly on antigens with or without a close pre-cutoff homolog.

Minjae Park, Roman Nett, Brian M. Petersen et al. · 0 citations
Open access Jul 2026

Analysing open-source protein folding models for nanobody binding prediction

These findings provide practical guidance for integrating open-source protein structure prediction models into AI-driven nanobody discovery pipelines while highlighting the need for improved generalization across antigens.

Yannick Vogt, Rebekka Roßberg, Jan Habermann et al. · 1 citation
Open access Aug 2026

Benchmarking Antibody Modeling Tools across Structure Prediction, Docking, and Paratope–Epitope Interface Analysis

This work evaluated ImmuneBuilder, IgFold, AlphaFold3, GRAMM, and dyMEAN on 50 non-redundant humanized antibody–antigen complexes using multiple retained predictions and paired statistical testing, finding all three antibody structure predictors were accurate.

Zeyuan Yu, Jilei Wu, Ziyao Ning et al. · 0 citations

A Hitchhiker’s Journey through Machine Learning for Structural Biology of Antibodies

This thesis examines the integration of machine learning into computational structural biology, with an emphasis on modelling and predicting antibody–antigen interactions. Such interactions are fundamental to numerous biological processes and are central to therapeutic antibody design. Despite recent advances in AI-based protein structure prediction, antibodies remain particularly challenging targets due to the high variability of their complementarity-determining regions, the limited availability of experimental structures, and the lack of strong co-evolutionary signal. To address these challenges, this work introduces several methodological contributions. In Chapter 2,DeepRank-GNN-esm incorporates embeddings from protein language models to replace computationally expensive evolutionary features, thereby improving both predictive performance and efficiency in scoring protein–protein complexes. In Chapter 3, a modelling pipeline is introduced that employs a flow-matching algorithm to effectively sample the conformational diversity of the antibody CDR-H3 loop. When integrated with ensemble docking, this approach significantly improves the accuracy of antibody–antigen complex modelling compared to existing methods. In Chapter 4, the thesis presents AbTune, a sequence-specific fine-tuning strategy for protein language models that enhances predictive performance across multiple antibody-related tasks, including structure prediction, mutation effect estimation, and binding affinity prediction, while remaining computationally efficient. In Chapter 5, DeepRank-Ab is developed as a geometric deep learning-based scoring function tailored to antibody–antigen complexes, achieving state-of-the-art performance in ranking near-native docking conformations. Chapter 6 summarizes the main findings of the thesis and discusses future research directions. Collectively, these contributions demonstrate how machine learning can be applied to address key limitations in antibody modelling and to facilitate the rational design of antibody-based therapeutics.

Xiaotong Xu · 0 citations
Book Open access Aug 2026

SAASBench: A Synthetic Antibody–antigen Specificity Benchmark

SAASBench provides a framework for evaluating the model's ability to estimate the specificity of a candidate antibody in relevant settings, indicating that strong performance on traditional affinity benchmarks does not automatically translate into reliable antibody specificity estimation in proteome-derived settings.

Dmitriy Umerenkov, Ivan Poddiakov · 0 citations
Open access Aug 2026

AVIDbase: A biologically accurate structural dataset of nanobody‐antigen complexes

Accurate structural data in a standardized format is one of the key factors behind the success of machine learning (ML)‐based methods for protein design and structure prediction. However, their application to nanobody‐antigen complexes has lower success rates compared to globular protein complexes, partly due to the limited amount of high‐quality structural data. While several dedicated databases already exist, automated assembly pipelines frequently overlook various artifacts, which act as additional noise and limit the effectiveness of ML applications. Common issues include incorrectly defined antigen assemblies, redundancy bias, inclusion of crystal contacts, strained geometry due to crystal packing, as well as missing density or post‐translational modifications near the interface. To address these issues, we present Antigen‐VHH Interface Database (AVIDbase), a highly curated dataset of nanobody‐antigen structures. In addition to correcting structural artifacts, the dataset provides a nonredundant set of structures with standardized chain identifiers, harmonized metadata, and cleaned atomic coordinates in a ready‐to‐use format for ML applications. AVIDbase is available on GitHub (github.com/Novartis/AVIDbase) and Zenodo (doi.org/10.5281/zenodo.20488703).

Tadej Medved, J. Lah, G. Miličić et al. · 0 citations