Skip to content
Open access

PepXPro: a framework for curating, generating, and optimizing structure-affinity protein-peptide datasets

Aug 2026 · bioRxiv · 0 citations · 34 references
Biology

TL;DR

P PepXPro is presented, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria that provides an extensible foundation for reproducible protein- peptide benchmark construction.

Abstract

Protein-peptide interactions are central to cellular signaling and to a growing class of peptide therapeutics, yet the datasets used to develop and benchmark computational methods for protein-peptide modeling remain poorly standardized. Available databases prioritize comprehensive coverage but require task-specific curation, while published benchmarks are typically distributed as static collections built with heterogeneous curation, quality-filtering, redundancy-reduction, and sampling strategies, limiting reproducibility and cross-study comparison. We present PepXPro, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria. PepXPro is organized into three components: Scrape, for deterministic curation of protein-peptide complex entries from public resources; GenSample, for constructing configurable subsets under explicit quality, redundancy, and sampling constraints; and Benchmark, for evaluating candidate subsets and selecting a nonredundant, representative, general-purpose benchmark for distribution. Starting from PDBbind and complementary resources, the curation pipeline yields a pool of proteinpeptide complex entries that retains chemically complex cases, including disulfide- linked cyclic peptides, which are commonly excluded from existing benchmarks. We release PepXPro Benchmark v1, a benchmark comprising 70 non-redundant protein- peptide complexes with experimentally determined structures and binding affinities. The underlying framework provides an extensible foundation for reproducible protein- peptide benchmark construction.

Read PDF

Similar papers

Open access Jul 2026

A comprehensive dataset of 32 million pentapeptide structures for high-throughput virtual screening.

Small peptides are widely used as binders, modulators, and structural motifs, but their conformational flexibility complicates structure-based analysis and high-throughput screening. We present an open dataset of three-dimensional structures for the complete space of canonical amino-acid pentapeptides: 3,200,000 unique sequences with up to 10 conformers per sequence, for a total of 32,000,000 peptide conformers. Structures were generated directly from sequence using an automated workflow built on UCSF ChimeraX for model construction, Reduce for hydrogen placement, and RDKit for conformer generation and optimization. The dataset is distributed as compressed archives with an accompanying index that maps each sequence and conformer identifier to its coordinate record, enabling efficient download, subset selection, and programmatic access. Technical validation includes symmetry-aware inter-conformer RMSD analysis, Ramachandran quality assessment, and benchmarking against experimentally observed pentapeptide fragments from the Protein Data Bank. Although we focus here on pentapeptides to enable exhaustive sequence coverage, the publicly released workflow is solely based on open-source software and can be applied to other short peptides to generate comparable conformer libraries. This resource supports virtual screening with pre-generated peptide conformer ensembles, method benchmarking, and machine-learning applications in peptide design and protein engineering by removing the need for researchers to repeatedly generate large conformer ensembles from scratch.

Josep-Ramon Codina, E. Dikici, Sapna K. Deo et al. · 0 citations
Open access Jul 2026

Minimal Data · Maximal Insight (MDMI): A Structure-guided Pipeline for Discovering Functional Alternatives in Peptide-Protein Interfaces

Minimal Data Maximal Insight (MDMI), a two-stage structure-guided computational pipeline that designs functional peptide variants using only a small, annotated dataset, demonstrates that structure-informed pipelines can uncover remote functional sequence space from minimal data.

P. Bayat, Spencer J. Perkins, Sebastian Clancy et al. · 0 citations
Open access Jul 2026

The Human Bindome: A Proteome-scale Atlas of Designed Binder Candidates

The Human Bindome is presented, a proteome-scale atlas of high-confidence in silico protein binder candidates that positions the Bindome as a resource of genetically encodable perturbagens for site-specific, modular control of protein function.

Julius Wenckstern, Anna M. Díaz-Rovira, Julia A. Kuhn et al. · 0 citations
Open access Jul 2026

Analysing open-source protein folding models for nanobody binding prediction

These findings provide practical guidance for integrating open-source protein structure prediction models into AI-driven nanobody discovery pipelines while highlighting the need for improved generalization across antigens.

Yannick Vogt, Rebekka Roßberg, Jan Habermann et al. · 0 citations
Open access Aug 2026

Profiling Protein‐Peptide Interactions by Yeast Surface Display

A protocol for discovering protein‐binding peptides using a very large, target‐agnostic yeast surface display library containing approximately 6.1 × 109 unique clones and providing broad coverage of short peptide sequence space is described.

Joseph D. Hurley, Andrew C. Kruse · 0 citations