Skip to content
Open access

Benchmarking Antibody Modeling Tools across Structure Prediction, Docking, and Paratope–Epitope Interface Analysis

Aug 2026 · Bioinformatics Advances · 0 citations

TL;DR

This work evaluated ImmuneBuilder, IgFold, AlphaFold3, GRAMM, and dyMEAN on 50 non-redundant humanized antibody–antigen complexes using multiple retained predictions and paired statistical testing, finding all three antibody structure predictors were accurate.

Abstract

Computational antibody engineering requires reliable prediction of antibody variable-fragment structures, antigen–antibody complexes, and binding interfaces. However, publicly available tools for these tasks have rarely been compared across the complete workflow under a controlled and statistically grounded design. We evaluated ImmuneBuilder, IgFold, AlphaFold3, GRAMM, and dyMEAN on 50 non-redundant humanized antibody–antigen complexes using multiple retained predictions and paired statistical testing. All three antibody structure predictors were accurate, with AlphaFold3 performing best overall and for the third complementarity-determining region of the heavy chain. AlphaFold3 also substantially outperformed GRAMM and dyMEAN in complex prediction, producing medium- or high-quality binding interfaces for 46% of the complexes, although overall interface accuracy remained limited. When docking was reliable, AlphaFold3 accurately recovered epitope and paratope residues, salt bridges, and non-bonded contacts, but reproduced hydrogen bonds and fine-grained contact strengths less consistently. These findings provide practical guidance for selecting tools across antibody-modeling workflows and identify persistent limitations in fine-grained interface prediction. Data, structural predictions, evaluation results, and analysis code are available from Zenodo under record 20710876.

Read PDF

Similar papers

Open access Aug 2026

Benchmarking antibody-antigen co-folding on human monomeric antigens

The most recent methods substantially outperformed earlier ones, producing medium-or-better top-ranked models for approximately half of post-cutoff Fv complexes without templates or experimental restraints, and performing similarly on antigens with or without a close pre-cutoff homolog.

Minjae Park, Roman Nett, Brian M. Petersen et al. · 0 citations
Open access Aug 2026

AFilter: Improved Antibody Epitope Prediction by Machine Learning-Optimized Interface Energy Filtering of AlphaFold3-Predicted Complex

A lightweight post-hoc filter that requires no re-docking and is directly compatible with existing AF3 prediction pipelines and transferable to other diffusion-based complex predictors, providing a practical quality-assurance layer for antibody epitope mapping in early-stage drug discovery.

Xiao-Yu Liu, Yu Wang · 0 citations
Open access Jul 2026

A systematic evaluation framework for universal antibody-antigen binding affinity prediction and candidate recommendation

This work proposes MochiBind, a sequence-only pairwise binding affinity predictor, and benchmark it against structure-derived baselines such as Boltz-2, GeoDock, and Graphinity, suggesting that sequence-based approaches can match or surpass structure-based models in generalization.

Yunrui Li, Yue Zhao, K. Sonmez et al. · 0 citations
Open access Jul 2026

Analysing open-source protein folding models for nanobody binding prediction

These findings provide practical guidance for integrating open-source protein structure prediction models into AI-driven nanobody discovery pipelines while highlighting the need for improved generalization across antigens.

Yannick Vogt, Rebekka Roßberg, Jan Habermann et al. · 1 citation

Does Surface Conservation Yield? Application to Data-Driven Docking

The interface prediction program WHISCY is presented, which combines surface conservation and structural information to predict protein–protein interfaces and demonstrates the potential of using interface predictions to drive protein–protein docking.

Sjoerd J. de Vries, A. V. van Dijk, A. M. Bonvin · 0 citations
Open access Jul 2026

Capabilities, specificity gaps and training-data dependence of AlphaFold3 across diverse application areas

It is found that, while AF3 can perform well in favourable settings, this performance is uneven across applications and its predictions and use of confidence metrics will depend strongly on the specific application area and must be interpreted with respect to training-set overlap.

O. Follonier, Yan Liu, Pablo Campomanes et al. · 1 citation