2026· Journal of Advances in Information Technology· 0 citations· 24 references
TL;DR
This study evaluates four encoding schemes—one-hot, k-mer (substring-based encoding), embeddings, and Position-Specific Scoring Matrix (PSSM) using Convolutional Neural Networks (CNNs) and Long Short-Term Memories (LSTMs) and shows that k-mer encoding achieved the highest accuracy.
Abstract
—In bioinformatics and computational biology, sequence classification is crucial for tasks such as protein function prediction, disease classification, and gene annotation. While deep learning has advanced this field, model performance is heavily influenced by the sequence encoding methods used. This study evaluates four encoding schemes—one-hot, k-mer (substring-based encoding), embeddings, and Position-Specific Scoring Matrix (PSSM) using Convolutional Neural Networks (CNNs) and Long Short-Term Memories (LSTMs). Annotated protein and DNA sequences were encoded, balanced, and trained under standardized conditions for fair comparison. Results show that k-mer encoding achieved the highest accuracy (89% with CNN, 90% with LSTM). LSTMs also performed well with embedding-based representations, effectively capturing sequence dependencies. In contrast, PSSM and one-hot encodings yielded lower accuracy, suggesting reduced suitability for deep learning. These findings provide practical guidance for selecting optimal model-encoding combinations, aiming to improve both accuracy and computational efficiency in sequence classification tasks.
Sequence-based protein druggability classification can support early target triage when structural information is unavailable, uncertain, or inconsistently linked to druggability labels. We present DrugPLMFormer, a sequence-first retrospective screening framework that combines frozen protein language model embeddings with self-attentive BiLSTM encoding, Transformer-based long-range modeling, optional physicochemical feature fusion, and compute-budgeted BO–CTCM model selection. Hyperparameters were selected through multi-fidelity screening within an approximately 200-evaluation budget, using a validation objective that combined AUPRC and MCC to balance threshold-free discrimination with operating-point stability, rather than to imply unrestricted generalization. On ProTar-II, using a 50% sequence-identity homology-aware split, DrugPLMFormer achieved 95.98% accuracy, 96.01% F1-score, 96.42% sensitivity, 95.61% specificity, and 0.981 ROC-AUC. Without using external data for training, tuning, threshold selection, or early stopping, the selected model showed favorable held-out mean performance on ProTar-II-Ind (96.62% accuracy, 0.9688 ROC-AUC) and DPI_CDF (96.20% accuracy, 0.9696 ROC-AUC). Paired external analyses indicated that accuracy and F1-score differences were numerically favorable but not statistically significant, whereas the ROC-AUC improvement on DPI_CDF was statistically supported. Train-to-external homology analysis showed that most external proteins had less than 50% sequence identity to the training set, although residual dataset shift and label heterogeneity may still affect generalization. With cached PLM embeddings, downstream CPU inference required approximately 1.0–1.2 ms per sequence, excluding tokenization and ESM-2 embedding generation. Overall, DrugPLMFormer provides a reproducible, leakage-aware framework for retrospective sequence-based druggability screening and target prioritization, while prospective validation and experimental confirmation remain necessary before operational deployment.
Z. Kafi, Khosro Rezaee, Hossein Eslami· Journal of King Saud Univers...· 0 citations
The novel combination of LZ complexity–based negative sample selection, CT feature representation, and GA-optimized CNN–LSTM architecture provides a robust and biologically informed framework for PPI prediction.
Predicting protein subcellular locations computationally is crucial for analyzing large protein datasets. A key issue is that similar sequences in training and test sets artificially inflate accuracy estimates. This study investigates whether Protein Language Model (PLM) features alone can achieve strong predictions using simple classifiers instead of complex architectures. We developed a streamlined deep learning framework combining pre-trained ESM-2 embeddings with an attention-enhanced Bi-LSTM network, deployed as a 3-fold ensemble with soft voting. Training used eukaryotic sequences with ≤40% similarity to ensure rigorous evaluation. The model achieved 86.81% accuracy (MCC = 0.825) on test data—a +22.27% improvement over an SVM baseline (64.54%, MCC = 0.530). On 86 newly released 2024 proteins, the system reached 88.37% accuracy (MCC = 0.827), surpassing DeepLoc 2.1 (80.23%, MCC = 0.714, p=0.007) and MULocDeep (77.91%, MCC = 0.682, p=0.019). However, the small validation set (N=86) and limited representation in categories like Mitochondrion (N=5) require cautious interpretation. The method only handles single-location assignments across four compartments, excluding multi-location proteins. Attention weight analysis shows the model identifies biologically relevant signals, including C-terminal membrane regions and N-terminal mitochondrial sequences, confirming that ESM-2 embeddings enable effective performance with simplified architectures.
Johaimen Omar· International journal of lif...· 0 citations
Antimicrobial resistance (AMR) has become a significant challenge in global public health. With the development of whole-genome sequencing, protein language models, graph neural networks, and multimodal learning, deep learning-based AMR prediction research has rapidly evolved from traditional sequence alignment and rule retrieval to representation learning and genotype-phenotype mapping for high-dimensional heterogeneous data. This review systematically summarizes the research progress in this field from three aspects: antimicrobial resistance gene (ARG) identification and classification, genomic mutation-driven resistance phenotype prediction, and non-WGS multimodal extension. The review shows that deep learning has significantly improved the modeling ability for distantly homologous sequences, complex mutation combinations, and heterogeneous data, driving AMR prediction from “database matching” to “learnable representations,” and from “single-label discrimination” to “multi-task, multi-representation, and multimodal fusion.” However, at the same time, problems such as dataset heterogeneity, inconsistent label standards, class imbalance, lineage mixing, insufficient external generalization, and insufficient interpretability still restrict the clinical application of these models. Future research should further strengthen the construction of basic models, standardized evaluation, mining of interpretable mechanisms, and joint modeling of multi-omics and clinical data to promote AMR prediction from method validation to real-world application.
The autoencoder framework encodes protein sequence information related to domains, families, and patterns—into a lengthy, sparse binary vector that outperforms other neural network models, including convolutional neural networks, recurrent neural networks, long short-term memory networks, and bidirectional long short-term memory networks.
Biswajit Senapati, Ranjita Das· Journal of Computer-Aided Mo...· 0 citations
CLDN18.2 is a promising tumor-specific antigen; however, the development of therapeutic antibodies against it is challenged by the need for simultaneous optimization of affinity and developability. To address this, we present cdrGPT, a deep learning framework based on GPT-2 for de novo generation of complementarity-determining region H3 (CDRH3) sequences. Our approach integrates pre-training on the Observed Antibody Space (OAS) database with structural templating derived from the known antibody zolbetuximab. Generated sequences were iteratively refined through rejection sampling and fine-tuned against a multi-parameter objective function encompassing predicted affinity and MHC class II binding risk. From an initial set of 50,000 sequences, this screening pipeline yielded 313 high-confidence candidates. Subsequent analysis using evolutionary scale modeling 2 (ESM2) embeddings, principal component analysis (PCA), and clustering revealed three structurally distinct clusters, with intra-cluster cosine similarities exceeding 0.99. Validation of seven representative sequences from the dominant cluster using AlphaFold3 confirmed high structural fidelity to the zolbetuximab template, demonstrating a root mean square deviation (RMSD) of 1.331 Å for the CDRH3 loop and positional deviations of less than 0.4 Å for key paratope residues. These results indicate that the designed variants preserve the core binding mode of the parent antibody. This study establishes a feasible pipeline for integrating AI-generated CDRH3 loops into functional antibody scaffolds, providing a foundation for the accelerated development of therapeutics targeting CLDN18.2 and other clinically relevant antigens.
Tao Qu, Lingyan Yuan, Weiran Cui et al.· PLoS Computational Biology· 0 citations