This study repurposed a machine learning algorithm to comprehensively chart the biosynthetic space of the biarylitides, including variation of precursor motifs, P450, and additional modifying enzymes, which yielded 277 biarylitide biosynthetic gene clusters (BGCs).
Abstract
Biarylitides are a group of bacterial ribosomally synthesized and post-translationally modified peptides (RiPPs) that contain a biaryl bridge formed by dedicated cytochrome P450 enzymes that can introduce different cross-links. The biarylitides are produced via a five-amino-acid precursor peptide, encoded by a minimal 18 bp gene that evades automatic detection. Previous genome mining approaches for biarylitides do not capture their full biosynthetic space. We therefore repurposed a machine learning algorithm to comprehensively chart the biosynthetic space of the biarylitides, including variation of precursor motifs, P450, and additional modifying enzymes, which yielded 277 biarylitide biosynthetic gene clusters (BGCs). We experimentally investigated biaryl formation with previously uninvestigated core peptide motifs, including YWH, YVH, and YWY, and elucidated the nature of these cross-links. This study significantly expands the biarylitide precursor and BGC diversity and provides directions for the systematic exploration of other RiPP families.
It is shown that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences, and coupling comparative genomics with protein machine learning can convert “small data” typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.
Xiaoyu Shan, I. Trindade, N. Glasser et al.· Proceedings of the National...· 0 citations
A systematic comparison of zero-shot ML models is provided and an iterative framework for integrating machine learning with directed evolution to accelerate enzyme engineering is established to accelerate enzyme engineering.
Daniel Gutierrez, Isa Madrigal Harrison, Aaron L. Feller et al.· bioRxiv· 0 citations
DeepAden achieves competitive performance compared with state-of-the-art tools on a benchmark dataset, and enabled the identification of two Streptomyces NRPS gene clusters through accurate A-domain substrates specificity predictions.
A small-sample, accelerated evolution strategy that integrates focused rational iterative site-specific mutagenesis (FRISM) with the EVOLVEpro model is reported, providing a robust, "lightweight" machine learning framework for the rapid development of new-to-nature photoenzymatic transformations.
Results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope, and suggest that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope.
Ravi G. Lal, Jason Yang, Ziyan Zhang et al.· bioRxiv· 0 citations