RulePep combines competitive predictive performance with additive logit reconstruction, rule-level evidence reporting, and predicate-suppression auditing, providing a transparent sequence-based framework for peptide candidate prioritization.
Abstract
Peptides are increasingly explored as therapeutic candidates, delivery vectors, and functional biomolecules, but experimental screening of peptide activity and safety remains costly because the sequence space is vast and small sequence changes can alter functionality. Computational peptide classification can therefore help prioritize candidates. However, many protein-language-model-based classifiers achieve strong performance using opaque prediction heads, making it difficult to determine which learned evidence supports or opposes a prediction. We present RulePep, an ESM-2-guided neural-symbolic classifier for peptide-function prediction. RulePep maps frozen ESM-2 sequence representation to learned latent predicates, polarity-constrained differentiable rules, and an additive symbolic logit whose components can be inspected at the case level. We evaluate RulePep on three biologically distinct peptide classification tasks: blood-brain barrier penetration, hemolytic potency, and anticancer activity. On the BBPpredict, HemoPI3, and AntiCP 2.0 alternate benchmark datasets, RulePep achieved AUROC/MCC values of 0.8869/0.6850, 0.9155/0.6820, and 0.9765/0.8633, respectively. Ablation experiments supported the contributions of multi-layer representation pooling, rule polarity, mined-rule initialization, symbolic capacity, and rule-derived aggregation. RulePep combines competitive predictive performance with additive logit reconstruction, rule-level evidence reporting, and predicate-suppression auditing, providing a transparent sequence-based framework for peptide candidate prioritization.
Sequence-based protein druggability classification can support early target triage when structural information is unavailable, uncertain, or inconsistently linked to druggability labels. We present DrugPLMFormer, a sequence-first retrospective screening framework that combines frozen protein language model embeddings with self-attentive BiLSTM encoding, Transformer-based long-range modeling, optional physicochemical feature fusion, and compute-budgeted BO–CTCM model selection. Hyperparameters were selected through multi-fidelity screening within an approximately 200-evaluation budget, using a validation objective that combined AUPRC and MCC to balance threshold-free discrimination with operating-point stability, rather than to imply unrestricted generalization. On ProTar-II, using a 50% sequence-identity homology-aware split, DrugPLMFormer achieved 95.98% accuracy, 96.01% F1-score, 96.42% sensitivity, 95.61% specificity, and 0.981 ROC-AUC. Without using external data for training, tuning, threshold selection, or early stopping, the selected model showed favorable held-out mean performance on ProTar-II-Ind (96.62% accuracy, 0.9688 ROC-AUC) and DPI_CDF (96.20% accuracy, 0.9696 ROC-AUC). Paired external analyses indicated that accuracy and F1-score differences were numerically favorable but not statistically significant, whereas the ROC-AUC improvement on DPI_CDF was statistically supported. Train-to-external homology analysis showed that most external proteins had less than 50% sequence identity to the training set, although residual dataset shift and label heterogeneity may still affect generalization. With cached PLM embeddings, downstream CPU inference required approximately 1.0–1.2 ms per sequence, excluding tokenization and ESM-2 embedding generation. Overall, DrugPLMFormer provides a reproducible, leakage-aware framework for retrospective sequence-based druggability screening and target prioritization, while prospective validation and experimental confirmation remain necessary before operational deployment.
Z. Kafi, Khosro Rezaee, Hossein Eslami· Journal of King Saud Univers...· 0 citations
It is shown that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification, and MPEP conditioning is established as a lightweight strategy for low-data peptide classification.
Joshua Almonte, M. Vu, Andrew Ahn et al.· bioRxiv· 0 citations
Bitter peptides are a practical barrier in food-grade protein hydrolysates, fermented products, and peptide-based supplements because they can compromise flavor before nutritional or functional value is realized. Sensory panels and mass-spectrometry-based identification remain reliable, but their throughput is limited for early screening of large peptide pools. Existing predictors usually emphasize either interpretable hand-crafted descriptors or deep sequence representations, whereas these two information sources may be complementary for food-oriented bitter peptide screening. Here, we propose iBitter-HF, a hybrid feature embedding method that integrates seven classes of hand-crafted descriptors with Unified Representation (UniRep) features. Light Gradient Boosting Machine (LGBM)-based feature-importance ranking was used to organize the candidate embeddings, and eXtreme Gradient Boosting (XGB) was used for classification of the selected feature subset. On the public BTP640 benchmark, the finalized 135-feature model achieved 96.9% accuracy on the independent test set. Literature-based comparison indicated competitive performance relative to eight reported bitter peptide predictors, and dimensionality reduction visualization suggested clearer local organization of bitter and non-bitter peptides after feature optimization. These results support iBitter-HF as a computational aid for sequence-level bitter peptide screening and debittering-oriented design of protein hydrolysates.
Feng Yan, Shi-Cheng Xiang, Yi Tang et al.· Foods· 0 citations
MolPACL is proposed, a prompt-augmentation-based supervised contrastive learning framework that incorporates high-level chemical semantics while preserving molecular identity, and achieves strong performance on both classification and regression tasks while reducing training cost.
Ali Forooghi, Luis Rueda, A. Ngom· IEEE transactions on computa...· 0 citations
Binary protein classification supports diverse tasks in computational biology, including pathway-membership inference and sequence-based candidate prioritization. Protein language models generate information-rich residue-level representations, but downstream classifiers commonly compress them using fixed pooling operations that may discard task-relevant sequence context. We present ETAP-CLF, a compact framework that combines pretrained per-residue ESM3 embeddings with lightweight transformer contextualization and learned attention pooling to classify variable-length proteins and generate residue-level attention scores. The ESM3 parameters remained frozen, and the same ETAP-CLF architecture and hyperparameter configuration were used across ferroptosis-, senescence-, and pyroptosis-associated protein prediction. ETAP-CLF achieved AUROCs of 0.98, 0.95 and 0.91 for these tasks, respectively. In the ferroptosis benchmark, ETAP-CLF outperformed the evaluated published models. These results demonstrate that a common downstream design can adapt to multiple process-associated classification tasks without fine-tuning the task-specific model architecture. ETAP-CLF provides a generalizable approach for sequence-based protein prioritization and a basis for broader evaluation across binary protein-classification problems.
Jianyu Ren, Hanli Jiang, Puchangxin Li et al.· bioRxiv· 0 citations
An unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets is presented and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.
Nazmul Shuzan, Jialun Wei, Jie Zheng· Journal of Chemical Informat...· 0 citations