P PepXPro is presented, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria that provides an extensible foundation for reproducible protein- peptide benchmark construction.
Abstract
Protein-peptide interactions are central to cellular signaling and to a growing class of peptide therapeutics, yet the datasets used to develop and benchmark computational methods for protein-peptide modeling remain poorly standardized. Available databases prioritize comprehensive coverage but require task-specific curation, while published benchmarks are typically distributed as static collections built with heterogeneous curation, quality-filtering, redundancy-reduction, and sampling strategies, limiting reproducibility and cross-study comparison. We present PepXPro, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria. PepXPro is organized into three components: Scrape, for deterministic curation of protein-peptide complex entries from public resources; GenSample, for constructing configurable subsets under explicit quality, redundancy, and sampling constraints; and Benchmark, for evaluating candidate subsets and selecting a nonredundant, representative, general-purpose benchmark for distribution. Starting from PDBbind and complementary resources, the curation pipeline yields a pool of proteinpeptide complex entries that retains chemically complex cases, including disulfide- linked cyclic peptides, which are commonly excluded from existing benchmarks. We release PepXPro Benchmark v1, a benchmark comprising 70 non-redundant protein- peptide complexes with experimentally determined structures and binding affinities. The underlying framework provides an extensible foundation for reproducible protein- peptide benchmark construction.
Small peptides are widely used as binders, modulators, and structural motifs, but their conformational flexibility complicates structure-based analysis and high-throughput screening. We present an open dataset of three-dimensional structures for the complete space of canonical amino-acid pentapeptides: 3,200,000 unique sequences with up to 10 conformers per sequence, for a total of 32,000,000 peptide conformers. Structures were generated directly from sequence using an automated workflow built on UCSF ChimeraX for model construction, Reduce for hydrogen placement, and RDKit for conformer generation and optimization. The dataset is distributed as compressed archives with an accompanying index that maps each sequence and conformer identifier to its coordinate record, enabling efficient download, subset selection, and programmatic access. Technical validation includes symmetry-aware inter-conformer RMSD analysis, Ramachandran quality assessment, and benchmarking against experimentally observed pentapeptide fragments from the Protein Data Bank. Although we focus here on pentapeptides to enable exhaustive sequence coverage, the publicly released workflow is solely based on open-source software and can be applied to other short peptides to generate comparable conformer libraries. This resource supports virtual screening with pre-generated peptide conformer ensembles, method benchmarking, and machine-learning applications in peptide design and protein engineering by removing the need for researchers to repeatedly generate large conformer ensembles from scratch.
Josep-Ramon Codina, E. Dikici, Sapna K. Deo et al.· Scientific Data· 0 citations
Minimal Data Maximal Insight (MDMI), a two-stage structure-guided computational pipeline that designs functional peptide variants using only a small, annotated dataset, demonstrates that structure-informed pipelines can uncover remote functional sequence space from minimal data.
P. Bayat, Spencer J. Perkins, Sebastian Clancy et al.· bioRxiv· 0 citations
The Human Bindome is presented, a proteome-scale atlas of high-confidence in silico protein binder candidates that positions the Bindome as a resource of genetically encodable perturbagens for site-specific, modular control of protein function.
Julius Wenckstern, Anna M. Díaz-Rovira, Julia A. Kuhn et al.· bioRxiv· 0 citations
These findings provide practical guidance for integrating open-source protein structure prediction models into AI-driven nanobody discovery pipelines while highlighting the need for improved generalization across antigens.
Yannick Vogt, Rebekka Roßberg, Jan Habermann et al.· Frontiers in Bioinformatics· 0 citations
A protocol for discovering protein‐binding peptides using a very large, target‐agnostic yeast surface display library containing approximately 6.1 × 109 unique clones and providing broad coverage of short peptide sequence space is described.
Joseph D. Hurley, Andrew C. Kruse· Current Protocols· 0 citations