AVIDbase: A biologically accurate structural dataset of nanobody‐antigen complexes
Abstract
Accurate structural data in a standardized format is one of the key factors behind the success of machine learning (ML)‐based methods for protein design and structure prediction. However, their application to nanobody‐antigen complexes has lower success rates compared to globular protein complexes, partly due to the limited amount of high‐quality structural data. While several dedicated databases already exist, automated assembly pipelines frequently overlook various artifacts, which act as additional noise and limit the effectiveness of ML applications. Common issues include incorrectly defined antigen assemblies, redundancy bias, inclusion of crystal contacts, strained geometry due to crystal packing, as well as missing density or post‐translational modifications near the interface. To address these issues, we present Antigen‐VHH Interface Database (AVIDbase), a highly curated dataset of nanobody‐antigen structures. In addition to correcting structural artifacts, the dataset provides a nonredundant set of structures with standardized chain identifiers, harmonized metadata, and cleaned atomic coordinates in a ready‐to‐use format for ML applications. AVIDbase is available on GitHub (github.com/Novartis/AVIDbase) and Zenodo (doi.org/10.5281/zenodo.20488703).