Skip to content
Open access

Accelerating protein engineering: an integrated framework combines protein language models and epistatic landscape modeling

Aug 2026 · Signal Transduction and Targeted Therapy · Vol 11 · 0 citations · 5 references
Medicine

TL;DR

MULTI-evolve is a model guided, universal, targeted installation of multimutants framework that rapidly designs hyperactive multimutant proteins and improves the identi fi cation of productive mutations compared with individual PLMs alone.

Abstract

In a recent study published in Science , Tran et al. introduced MULTI-evolve (model guided, universal, targeted installation of multimutants), a machine learning-guided directed evolution (MLDE) framework that rapidly designs hyperactive multimutant proteins. 1 By integrating protein language models (PLMs), epistasis-aware modeling, and a high-ef fi ciency multisite mutagenesis platform, the authors address the major combinatorial challenge that has long bottlenecked protein engineering and therapeutic development. Protein function is encoded by amino acid sequence, yet discovering productive mutational combinations remains dif fi cult because protein fi tness landscapes are high-dimensional and strongly shaped by epistatic interactions. Directed evolution has long provided a powerful route for protein optimization through repeated mutagenesis and screening. 2 More recently, MLDE has expanded accessible sequence space through in silico screening, but its performance still depends on accurately capturing epistatic interactions, the context-dependent effects between mutations that shape protein fi tness. 3 In particular, many existing approaches require large training datasets, multiple experimental rounds, or labor-intensive synthesis, and remain limited in extrapolating to high-order multimutants because of incomplete epistatic modeling. Tran et al. address these limitations with an end-to-end framework for direct multimutant exploration. MULTI-evolve is built on three conceptual components. First, the framework deploys a PLM zero-shot ensemble approach to nominate function-enhancing single mutations. By combining structure-informed and sequence-based models with z-score normalization, the authors improved the identi fi cation of productive mutations compared with individual PLMs alone. Second, they trained fully connected neural networks (FCNNs) on a compact dataset of experimentally characterized single and double mutants to learn epistatic interactions and predict higher-order variants. This data-ef

Read PDF

Similar papers

Jul 2026

Prioritizing Stability-enhancing Mutations using the ESM Protein Language Model in conjunction with Physics-based MM/GBSA Predictions.

This work investigates the effectiveness of two different computational methods as triaging tools for prioritizing target positions and identifying specific mutations that are likely to improve protein thermodynamic stability and proposes a hybrid mutation prioritization and selection strategy that achieves better accuracy than either method alone.

Emily R. Rhodes, G. Scarabelli, Jonathan Jou et al. · 0 citations
Open access Aug 2026

Evolution-inspired multi-objective Bayesian optimization for protein engineering

EvoMOBO is established as a modular framework for multi-objective protein engineering using experimental or mechanism-derived labels using simulation-derived mechanistic descriptors, with experiments reserved for final validation.

Kai Wen, Sirui Wang, Yixin Sun et al. · 0 citations
#protein folding Open access Aug 2026

Adaptive model-guided protein evolution with sparse data optimizes compact eukaryotic genome editors.

EvoMax is established as an integrated strategy for engineering compact eukaryotic Fz2 genome editors and FanzMAX v3-hLa is identified as a high-efficiency programmable nuclease for mammalian genome editing.

Shijie Wan, Jackson Gold, Pranay Vure et al. · 0 citations
Open access Aug 2026

Directed Evolution in Codon Space

Directed evolution is commonly used in protein engineering, where mature molecules are routinely improved through iterative local search of amino acid space. Here, we extend this principle to coding DNA. We developed a language-model-guided framework that iteratively refined industry-optimized coding sequences of clinical-stage therapeutics through synonymous exploration of codon space. Across 23 antibody-based therapeutics, SynCodonLM-guided refinement significantly increased recombinant expression in CHO cells for 17 molecules (74% responder rate), without significant compromise of product-quality or biophysical attributes. Moreover, changes in model likelihood predicted expression gains more effectively than heuristic statistical or mRNA-structure descriptors, despite no explicit expression objective. Codon-level likelihood also tracked temporal progression in influenza A H1N1 sequences, indicating the model captures evolutionary signal. These results show that even production-optimized sequences retain accessible fitness in synonymous codon space, establishing directed evolution as a practical strategy to improve biologic expression, a key manufacturing bottleneck, without altering protein sequence.

James Heuschkel, Laura Kingsley, Jon Reed et al. · 0 citations
Open access Aug 2026

Coevolution-informed Bayesian optimization for sample-efficient protein design

This work introduces ALSEBO (Active Learning Sequence Exploration via Bayesian Optimization), which couples a generative latent sequence landscape to Bayesian optimization and featurizes candidates with direct-coupling-analysis (DCA) coevolutionary statistics.

D. P. Kulathunga, Divyanshu Shukla, D. Potoyan · 0 citations
Open access Jul 2026

Miniaturizing and modifying natural proteins with Raygun.

Proteins have evolved over billions of years through coordinated substitutions, insertions and deletions, yet computational protein design cannot fully replicate nature's ability to engineer new proteins from existing templates. Protein language models1-3 generate informative per-residue representations, but harnessing them for large-scale, function-preserving sequence modifications has remained beyond reach. Here we introduce Raygun, a generative artificial intelligence framework that enables miniaturization, modification and augmentation of proteins, using a probabilistic encoding of protein sequences constructed from language model embeddings. Our key conceptual advance is to encode each protein not as a sequence of variable length in high-dimensional space, but as a probability distribution in fixed dimensions, making proteins of any length directly commensurable. Controlled by just two parameters governing substitutions and length changes, Raygun can shrink proteins by 10-25% (sometimes more than 50%), expand them beyond their natural size, and introduce extensive sequence diversity, all while preserving predicted structural integrity and functional sites. In cell-based validation, Raygun miniaturized fluorescent proteins (2 shorter than 96% of fluorescent proteins in FPbase) and TurboID, a synthetic biotin ligase that has been widely adopted for proteomics. It also expanded epidermal growth factor (EGF), generating variants with higher EGFR-binding affinity than the wild type. These results show that protein function can be faithfully captured in a length-agnostic representation, enabling the kind of coordinated, large-scale sequence modifications that characterize natural protein evolution.

Kapil Devkota, Daichi Shonai, Joey Mao et al. · 1 citation