Aug 2026· Journal of Computer-Aided Molecular Design· Vol 40· 0 citations· 37 references
Medicine
TL;DR
The autoencoder framework encodes protein sequence information related to domains, families, and patterns—into a lengthy, sparse binary vector that outperforms other neural network models, including convolutional neural networks, recurrent neural networks, long short-term memory networks, and bidirectional long short-term memory networks.
This mini review traces the evolution of AI-driven methods in protein research, from early residue-contact prediction using coevolutionary information to transformative breakthroughs, the rise of protein language models (PLMs), and the emerging era of generative design and functional modeling.
Guodong Min, Huan Peng· Methods in molecular biology· 0 citations
A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.
Carl David Jasper Causin, M. Fyta· APL Machine Learning· 0 citations
The novel combination of LZ complexity–based negative sample selection, CT feature representation, and GA-optimized CNN–LSTM architecture provides a robust and biologically informed framework for PPI prediction.
The Human Genome Project (HGP) was a large international research effort that timelined between 1990 and 2003, marking the successful mapping of the entire human genome. Despite the promising performance of deep learning models, especially LLMs like Bio-BERT and ProtBERT on well-annotated datasets. Bio-BERT is a pre-trained language model designed for biomedical corpora, it is not able to independently predict gene functionality. A pre-trained model called ProtBERT uses protein sequences to learn features. It is unable to predict gene functionality directly, without more training data. In order to overcome these challenges, a Gene Bio-BERT based framework is proposed for automated gene function prediction utilizing deep learning methods in biomedical data analysis. This Gene Bio-BERT Framework is divided into 3 modules such as Data collection and Preprocessing, Gene Bio-BERT model training, Feature aggregation layer and prediction of functionality. The initial module focuses on data collection and preprocessing. Data is collected using entrez API from NCBI (National Center for Biotechnology Information) to retrieve human gene data. Preprocessing techniques like Tokenization and feature extraction are then used to handle the data. In the second module, a Gene Bio-BERT transformer encoder with an attention-based feature fusion layer and optimized hyperparameters is used to train the model and learn contextual embeddings. The third module generates results by using aggregated transformer representations to produce functional predictions for unknown genes. In predicting gene function, the proposed Gene Bio-BERT model attains an exceptional accuracy of 94.5% and F1 score of 0.87. Additionally, the model's predictive accuracy remains similar when tested on an unannotated gene which gives similarity score 0.84.
A long-context protein language model is introduced, trained on entire viral genomes, leveraging a biologically informed sparse attention mechanism in which inter-protein links are inferred computationally and used as sparsity priors.
T. Dejean, Barbra D. Ferrell, Zachary D. Schreiber et al.· GigaScience· 0 citations
The utility of HA sites for suggesting candidate binding sites and the biological interpretability of PLM representations is explored, demonstrating the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.
Sophia J. Pribus, Russ B. Altman, Gowri Nayar· bioRxiv· 0 citations