Skip to content
#protein folding Open access

Quantifying the Provenance-to-Function Gap in Antidiabetic Peptide Prediction: Homology-Aware Evaluation and the ADP-Hybrid Baseline

Oct 2026 · bioRxiv · 0 citations · 58 references
Biology

TL;DR

Auditing sixteen further peptide benchmarks shows the coupling is not confined to this resource: fragment families share a label more often than chance in eleven of fourteen testable datasets, and not in three, so the property is common rather than universal.

Abstract

Antidiabetic peptides (ADPs) are short bioactive sequences of therapeutic interest, and sequence-based classifiers prioritise experimental candidates. A classifier is useful only if its accuracy transfers to unseen sequences, which depends on how the benchmark was assembled and partitioned. We term the distance between what such a classifier is scored on and the function it is meant to predict the provenance-to-function gap, and we quantify it. We re-evaluate the two-layer ADP benchmark of Basith et al. under a protocol that groups homologous peptides rather than splitting them at random. Of the 877 ADPs, 218 attribute to a precursor protein by exact substring containment; within that subset the second-layer label coincides exactly with precursor identity, all 140 human-insulin fragments carrying the type-1 label and all 76 bovine milk-protein fragments the type-2 label, without exception. Accuracy tracks identity to the training set, rising from a Matthews correlation coefficient (MCC) of 0.39–0.43 below 50% identity to 0.92–0.96 between 70% and 90%, and peptide length alone reaches MCC 0.619 on held-out data. Auditing sixteen further peptide benchmarks shows the coupling is not confined to this resource: fragment families share a label more often than chance in eleven of fourteen testable datasets, and not in three, so the property is common rather than universal. Rebuilding the published architecture on identical folds shows its advantage over a single classifier is a function of the split: present under random partitioning, absent once homologues are separated. ADP-Hybrid, one tree-ensemble classifier per layer, reaches MCC 0.843 on layer 1 against a published 0.841 at 21–38 times the inference throughput, and 0.801 against 0.858 on layer 2. We release the protocol and the controls that expose these properties.

Read PDF

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or seque...

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.