Aug 2026· ACS Catalysis· Vol 16, pp. 16076-16091· 0 citations· 55 references
TL;DR
Overall, MmALS represents a data-efficient active learning framework for protein evolution, advancing computational enzyme design by enabling accurate optimization under limited experimental sampling.
Abstract
Protein language models have shown strong potential in modeling fitness landscapes for directed evolution; however, their predictive accuracy and generalization to unexplored sequence space remain limited under few-shot learning conditions. Here, we present MmALS (Multi-modal Active Learning System), a few-shot active learning framework that incorporates cold-start region scanning and multi-objective optimization to enable unbiased mutation-site selection. Its learning module employs a multimodal fusion architecture that incorporates sequence, structural, and multiple sequence alignment information, enabling accurate variant fitness prediction under low-sample conditions. The robustness of the MmALS logical framework was validated through in silico benchmarking across 12 independent datasets, in which it consistently accelerated convergence and demonstrated transferability across diverse enzyme fitness landscapes. Using Caldicellulosiruptor saccharolyticus Cellobiose 2-epimerase (CsCE) as the experimental validation, MmALS achieved a 3.88-fold enhancement of isomerization activity after only three iterative rounds (approximately 50 mutants per round), reaching a record-high activity of 20.72 U/mg. Overall, MmALS represents a data-efficient active learning framework for protein evolution, advancing computational enzyme design by enabling accurate optimization under limited experimental sampling.
EvoMOBO is established as a modular framework for multi-objective protein engineering using experimental or mechanism-derived labels using simulation-derived mechanistic descriptors, with experiments reserved for final validation.
Kai Wen, Sirui Wang, Yixin Sun et al.· bioRxiv· 0 citations
Three modeling frameworks are developed, including models based on handcrafted features, models using embedding representations extracted from ProteinMPNN, and ensemble models integrating a diverse set of state‐of‐the‐art predictors integrating a diverse set of state‐of‐the‐art predictors.
Yang Liu, Jian Zhang, Minghui Li· Protein Science· 0 citations
CLDN18.2 is a promising tumor-specific antigen; however, the development of therapeutic antibodies against it is challenged by the need for simultaneous optimization of affinity and developability. To address this, we present cdrGPT, a deep learning framework based on GPT-2 for de novo generation of complementarity-determining region H3 (CDRH3) sequences. Our approach integrates pre-training on the Observed Antibody Space (OAS) database with structural templating derived from the known antibody zolbetuximab. Generated sequences were iteratively refined through rejection sampling and fine-tuned against a multi-parameter objective function encompassing predicted affinity and MHC class II binding risk. From an initial set of 50,000 sequences, this screening pipeline yielded 313 high-confidence candidates. Subsequent analysis using evolutionary scale modeling 2 (ESM2) embeddings, principal component analysis (PCA), and clustering revealed three structurally distinct clusters, with intra-cluster cosine similarities exceeding 0.99. Validation of seven representative sequences from the dominant cluster using AlphaFold3 confirmed high structural fidelity to the zolbetuximab template, demonstrating a root mean square deviation (RMSD) of 1.331 Å for the CDRH3 loop and positional deviations of less than 0.4 Å for key paratope residues. These results indicate that the designed variants preserve the core binding mode of the parent antibody. This study establishes a feasible pipeline for integrating AI-generated CDRH3 loops into functional antibody scaffolds, providing a foundation for the accelerated development of therapeutics targeting CLDN18.2 and other clinically relevant antigens.
Tao Qu, Lingyan Yuan, Weiran Cui et al.· PLoS Computational Biology· 0 citations
This study evaluates four encoding schemes—one-hot, k-mer (substring-based encoding), embeddings, and Position-Specific Scoring Matrix (PSSM) using Convolutional Neural Networks (CNNs) and Long Short-Term Memories (LSTMs) and shows that k-mer encoding achieved the highest accuracy.
T. Kurniawan, Deshinta Arova Dewi, Randy Joy Magno Ventayen· Journal of Advances in Infor...· 0 citations
Cutting-edge bioinformatics research is increasingly intertwined with pre-trained model techniques. However, achieving superior performance of these models in downstream applications typically requires large amounts of accurately labeled experimental data for fine-tuning, which poses substantial practical challenges due to the difficulty in preparing such datasets at scale. To address this limitation, we propose a novel few-shot fine-tuning framework, the Evolution-Aware Adaptation (Evo-AA). It aligns fine-tuning with pre-training objectives while integrating prompt learning and biological coevolutionary insights. Additionally, we introduce reinforced prompting and lambda ranking loss to further improve performance. Extensive experiments demonstrate that Evo-AA with limited training set, enhances the spearman correlation in fine-tuning tasks, while achieving superior precision and recall rates in homology search tasks. Our findings suggest that Evo-AA holds great potential to drive advancements in protein engineering and computational biology.
Yuxuan Wu, Huiqun Yu, Guisheng Fan et al.· Annual International Compute...· 0 citations
Sequence-based protein druggability classification can support early target triage when structural information is unavailable, uncertain, or inconsistently linked to druggability labels. We present DrugPLMFormer, a sequence-first retrospective screening framework that combines frozen protein language model embeddings with self-attentive BiLSTM encoding, Transformer-based long-range modeling, optional physicochemical feature fusion, and compute-budgeted BO–CTCM model selection. Hyperparameters were selected through multi-fidelity screening within an approximately 200-evaluation budget, using a validation objective that combined AUPRC and MCC to balance threshold-free discrimination with operating-point stability, rather than to imply unrestricted generalization. On ProTar-II, using a 50% sequence-identity homology-aware split, DrugPLMFormer achieved 95.98% accuracy, 96.01% F1-score, 96.42% sensitivity, 95.61% specificity, and 0.981 ROC-AUC. Without using external data for training, tuning, threshold selection, or early stopping, the selected model showed favorable held-out mean performance on ProTar-II-Ind (96.62% accuracy, 0.9688 ROC-AUC) and DPI_CDF (96.20% accuracy, 0.9696 ROC-AUC). Paired external analyses indicated that accuracy and F1-score differences were numerically favorable but not statistically significant, whereas the ROC-AUC improvement on DPI_CDF was statistically supported. Train-to-external homology analysis showed that most external proteins had less than 50% sequence identity to the training set, although residual dataset shift and label heterogeneity may still affect generalization. With cached PLM embeddings, downstream CPU inference required approximately 1.0–1.2 ms per sequence, excluding tokenization and ESM-2 embedding generation. Overall, DrugPLMFormer provides a reproducible, leakage-aware framework for retrospective sequence-based druggability screening and target prioritization, while prospective validation and experimental confirmation remain necessary before operational deployment.
Z. Kafi, Khosro Rezaee, Hossein Eslami· Journal of King Saud Univers...· 0 citations