Jul 2026· Annual International Computer Software and Applications Conference· pp. 1-10· 0 citations· 34 references
Abstract
Cutting-edge bioinformatics research is increasingly intertwined with pre-trained model techniques. However, achieving superior performance of these models in downstream applications typically requires large amounts of accurately labeled experimental data for fine-tuning, which poses substantial practical challenges due to the difficulty in preparing such datasets at scale. To address this limitation, we propose a novel few-shot fine-tuning framework, the Evolution-Aware Adaptation (Evo-AA). It aligns fine-tuning with pre-training objectives while integrating prompt learning and biological coevolutionary insights. Additionally, we introduce reinforced prompting and lambda ranking loss to further improve performance. Extensive experiments demonstrate that Evo-AA with limited training set, enhances the spearman correlation in fine-tuning tasks, while achieving superior precision and recall rates in homology search tasks. Our findings suggest that Evo-AA holds great potential to drive advancements in protein engineering and computational biology.
This work demonstrates how to provide task-specific information without losing the general knowledge learned during pretraining by using direct preference optimization to align a structure-conditioned protein language model to preferentially generate stable protein sequences.
Talal Widatalla, Ashir Borah, Samuel H. King et al.· Nature Methods· 0 citations
Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.
Ying Du, Zhikang Liu, Jing Wan· Journal of Microbiological M...· 0 citations
Protein language models (pLMs) offer great potential for protein sequence analysis, yet the scarcity of labeled data often limits their effectiveness in fine-tuning. Data augmentation is a promising remedy, but systematic evaluation of augmentation strategies for protein sequences remains limited, and the conditions under which augmentation confers downstream benefits are not well understood. In this paper, we systematically investigate pLM-guided substitution-based augmentation across seven protein prediction tasks. We propose ProtAug, a framework that leverages encoder-based (ESM-2) and autoregressive (ProtGPT2) pLMs to generate augmented sequences with user-controlled variation levels. Our investigation focuses on four questions: (Q1) whether pLM-synthesized sequences preserve more original signals than simpler methods, (Q2) to what extent augmentation improves prediction performance, (Q3) how variation levels affect downstream accuracy across tasks and models, and (Q4) whether biological plausibility is a necessary condition for achieving improvement. Our experimental results show that: (1) ProtAug Esm generally preserves motifs and structural similarity better than simple substitution, often comparable to homology retrieval; (2) augmentation yields consistent but task-dependent improvements, with ProtAug Esm achieving the best or second-best performance in 5 out of 7 tasks at 10% variation; (3) low-to-moderate variation levels (2–30%) perform best overall, although high-variation augmentation can benefit certain structure-related tasks; (4) the necessity of biological plausibility is task- and variation-dependent—while semantic preservation correlates with performance at low-to-moderate variation levels, improved generalization at high variation levels suggests that regularization effects, rather than label preservation, can also drive performance gains.
Predicting protein stability, like changes in melting temperature (ΔTm) caused by mutations, is a critical task in therapeutic protein engineering and drug discovery. This is reflected by a growing solution space, including both AI-based sequence and structure based methods. This paper demonstrates that accurate ΔTm prediction does not require structural input features, but can achieve state-of-the-art results with a careful training design for large sequence-based protein language models. We combine an autoresearch-inspired setup search with controlled ablation studies and show that a well-tuned sequence-only ESM2-650M model [6] outperforms structure-informed methods in our benchmark, achieving the lowest error (MAE/RMSE) and competitive Pearson correlation without pH or structural inputs. We further show that choices such as loss function, pooling strategy, auxiliary supervision, and finetuning regime materially affect performance.
Daniel Siegismund, Mario Wieser, E. Natali et al.· bioRxiv· 0 citations
The utility of HA sites for suggesting candidate binding sites and the biological interpretability of PLM representations is explored, demonstrating the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.
Sophia J. Pribus, Russ B. Altman, Gowri Nayar· bioRxiv· 0 citations
It is found that many GigaRef singletons belong to a cluster under alternative parameter settings, suggesting that genomic and metagenomic datasets may require dataset-specific clustering configurations, and it is shown that singletons share mutual information with clustered sequences, making them learnable by PLMs and useful for training.
R. Vinod, Samir Char, Ava A. Amini et al.· bioRxiv· 0 citations