Aug 2026· TIBS -Trends in Biochemical Sciences. Regular ed· 0 citations· 68 references
Medicine
TL;DR
How solid-phase peptide synthesis, genetically encoded libraries, and high-throughput selection and screening enable systematic exploration of vast, noncanonical landscapes largely inaccessible to traditional engineering is discussed.
Abstract
Research on the structural and functional potential of minimal amino acid alphabets is shifting from reductive 'pruning' of extant proteins to bottom-up exploration of combinatorial sequence space. This review highlights the experimental and computational toolkits driving this transition. We discuss how solid-phase peptide synthesis, genetically encoded libraries, and high-throughput selection and screening enable systematic exploration of vast, noncanonical landscapes largely inaccessible to traditional engineering. Integrating these approaches with generative machine learning and de novo protein design allows researchers to move beyond observing what evolution produced toward exploring what chemistry permits. Together, these advances are redirecting the field from simplifying extant proteins to uncovering the physicochemical principles governing protein foldability, while expanding the design space beyond the constraints imposed by biological evolution.
Proteins have evolved over billions of years through coordinated substitutions, insertions and deletions, yet computational protein design cannot fully replicate nature's ability to engineer new proteins from existing templates. Protein language models1-3 generate informative per-residue representations, but harnessing them for large-scale, function-preserving sequence modifications has remained beyond reach. Here we introduce Raygun, a generative artificial intelligence framework that enables miniaturization, modification and augmentation of proteins, using a probabilistic encoding of protein sequences constructed from language model embeddings. Our key conceptual advance is to encode each protein not as a sequence of variable length in high-dimensional space, but as a probability distribution in fixed dimensions, making proteins of any length directly commensurable. Controlled by just two parameters governing substitutions and length changes, Raygun can shrink proteins by 10-25% (sometimes more than 50%), expand them beyond their natural size, and introduce extensive sequence diversity, all while preserving predicted structural integrity and functional sites. In cell-based validation, Raygun miniaturized fluorescent proteins (2 shorter than 96% of fluorescent proteins in FPbase) and TurboID, a synthetic biotin ligase that has been widely adopted for proteomics. It also expanded epidermal growth factor (EGF), generating variants with higher EGFR-binding affinity than the wild type. These results show that protein function can be faithfully captured in a length-agnostic representation, enabling the kind of coordinated, large-scale sequence modifications that characterize natural protein evolution.
Kapil Devkota, Daichi Shonai, Joey Mao et al.· Nature· 1 citation
This mini review traces the evolution of AI-driven methods in protein research, from early residue-contact prediction using coevolutionary information to transformative breakthroughs, the rise of protein language models (PLMs), and the emerging era of generative design and functional modeling.
Guodong Min, Huan Peng· Methods in molecular biology· 0 citations
Biomolecular condensates form through phase separation, playing a crucial role in cellular organization and gene regulation. However, native condensate elements often suffer from high molecular weight and sequence redundancy, limiting their use in prokaryotic systems. To address this, we established an integrated framework combining generative artificial intelligence, computational prediction, and experimental validation. We began by constructing a benchmark dataset of eukaryotic-derived phase-separation elements in Escherichia coli. This dataset was used to build the PSVAE model for de novo sequence design and the LMPsPred classifier for high-throughput screening. Finally, we identified 13 short-sequence elements with low molecular weight (16-26 kDa) and minimal redundancy. Experimental validation confirmed that these elements formed condensates in E. coli, with enhanced protein recruitment compared to native elements. This study highlights the potential of AI-guided design to generate condensate-like assemblies and expands the repertoire of candidate phase-separation elements for prokaryotic systems.
Di Zhang, Bin-Yun Zhang, Cheng Zheng et al.· Synthetic and Systems Biotec...· 0 citations
A protein design strategy is used that couples a structure-guided inverse-folding model with evolution-informed residue constraints to generate active, divergent variants of TnpB, a minimal CRISPR-Cas12-like nuclease, termed SynTnpBs, establishing a strategy for creating non-natural RNA-guided nucleases and conformationally active nucleic acid binders, enlarging the designable protein space.
Petr Skopintsev, Isabel Esain-Garcia, Evan C. DeTurk et al.· Science· 2 citations
This review examines current computational strategies for exploring constrained protein fitness landscapes, including sequence-derived evolutionary descriptors, structural fitness assessment, energetic evaluation, and integrated multi-parameter scoring.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.