iBitter-HF: A Method for Bitter Peptide Sequence Identification Based on Hybrid Feature Embedding
Abstract
Bitter peptides are a practical barrier in food-grade protein hydrolysates, fermented products, and peptide-based supplements because they can compromise flavor before nutritional or functional value is realized. Sensory panels and mass-spectrometry-based identification remain reliable, but their throughput is limited for early screening of large peptide pools. Existing predictors usually emphasize either interpretable hand-crafted descriptors or deep sequence representations, whereas these two information sources may be complementary for food-oriented bitter peptide screening. Here, we propose iBitter-HF, a hybrid feature embedding method that integrates seven classes of hand-crafted descriptors with Unified Representation (UniRep) features. Light Gradient Boosting Machine (LGBM)-based feature-importance ranking was used to organize the candidate embeddings, and eXtreme Gradient Boosting (XGB) was used for classification of the selected feature subset. On the public BTP640 benchmark, the finalized 135-feature model achieved 96.9% accuracy on the independent test set. Literature-based comparison indicated competitive performance relative to eight reported bitter peptide predictors, and dimensionality reduction visualization suggested clearer local organization of bitter and non-bitter peptides after feature optimization. These results support iBitter-HF as a computational aid for sequence-level bitter peptide screening and debittering-oriented design of protein hydrolysates.