Jul 2026· Proceedings of the National Academy of Sciences of the United States of America· Vol 123, pp. e2607571123 - e2607571123· 0 citations· 51 references
Medicine
TL;DR
It is shown that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences, and coupling comparative genomics with protein machine learning can convert “small data” typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.
Abstract
Machine learning has enabled powerful biological discoveries using models trained on large datasets. However, for many important biological questions, such as identifying enzymes that transform understudied substrates, sparsity of training data is often a major bottleneck. Here, using phenazine natural products as a case study, we show that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences. We name this approach ML-CITO (Machine Learning for genomic Context-Informed Transferable discOvery). Applying this framework led to the discovery of PTC (Phenazine-Thiol Conjugase), the first enzyme known to catalyze phenazine thioconjugation, a phenazine modification reaction long observed but previously presumed to occur only through non-enzymatic chemistry. In silico simulation and experimental measurements demonstrate that PTC binds to both phenazine and glutathione as substrates. Recombinant expression and biochemical characterization reveal that PTC promotes glutathione-dependent modification of phenazines, yielding distinct reaction outcomes that depend on substrate identity. Although thiol-conjugated phenazine products exhibit reduced toxicity to bacterial cells, deletion of the gene encoding PTC does not confer a strong fitness disadvantage, illustrating how direct learning of sequences can uncover relevant enzymes that might evade phenotype-based genetic screens. Together, these results demonstrate that coupling comparative genomics with protein machine learning can convert “small data” typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.
This study repurposed a machine learning algorithm to comprehensively chart the biosynthetic space of the biarylitides, including variation of precursor motifs, P450, and additional modifying enzymes, which yielded 277 biarylitide biosynthetic gene clusters (BGCs).
Leo Padva, Jemma Gullick, Friederike Biermann et al.· JACS Au· 0 citations
Results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope, and suggest that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope.
Ravi G. Lal, Jason Yang, Ziyan Zhang et al.· bioRxiv· 0 citations
PubCheF-1, a deep learning model that predicts literature-derived biological function directly from chemical structure, establishes that machine learning-based prediction of biological function derived from the language of scientific literature allows the identification of bioactive molecules at high hit rates, thereby accelerating therapeutic discovery.
Clayton W. Kosonocky, Nikol Kadeřábková, Kangsan Kim et al.· bioRxiv· 0 citations
These findings introduce H_1 as a computationally prioritized, putative MLK4-binding lead and provide a hypothesis-generating framework for MLK4-targeted scaffold prioritization, while recognizing that experimental activity and kinome selectivity profiling remain necessary before H_1 can be described as a confirmed MLK4 inhibitor or MLK4-selective compound.
Afnan A. Alzaghari, S. Daoud, Husam Nassar et al.· Journal of Pharmaceutical In...· 0 citations
Four recently released ES and ER prediction models are benchmarked and it is suggested that interaction-aware representations from full biomolecular complexes may provide a promising basis for enzyme prioritization.
Elizabeth H. Mahood, N. Komorníková, Tom'avs Pluskal et al.· 0 citations
A systematic comparison of zero-shot ML models is provided and an iterative framework for integrating machine learning with directed evolution to accelerate enzyme engineering is established to accelerate enzyme engineering.
Daniel Gutierrez, Isa Madrigal Harrison, Aaron L. Feller et al.· bioRxiv· 0 citations