Aug 2026· Synthetic and Systems Biotechnology· Vol 17, pp. 133 - 143· 0 citations· 48 references
Medicine
TL;DR
This work establishes the first deep learning-enabled RBS design platform for P. denitrificans, offering a robust tool for precise translational regulation and advancing synthetic biology applications in environmental biotechnology.
Abstract
Paracoccus denitrificans is a heterotrophic nitrifying–aerobic denitrifying bacterium widely used in wastewater treatment and as a model system for studying electron transport chains, making it a promising chassis for environmental synthetic biology. Precise translational control of gene expression is essential for engineering desired traits in this organism, yet the design of functional ribosome binding sites (RBS) in P. denitrificans remains hindered by limited understanding of their sequence–activity relationships. To address this gap, we systematically profiled 1335 native RBS sequences via integrated transcriptomic and proteomic analyses. We found that RBS strength is predominantly governed by a purine-rich Shine–Dalgarno motif located 5–8 bp upstream of the start codon, with specific A/G patterns in this region serving as key determinants. Building on this dataset, we developed and compared CNN, BiLSTM, and Transformer models for RBS strength prediction; among them, the CNN achieved the highest predictive correlation (Pearson r = 0.68). Additionally, a WGAN-GP framework was implemented to generate novel RBS sequences, which were subsequently filtered and evaluated by the prediction framework to enable the design of RBSs with user-specified strengths. Experimental validation showed a relatively strong correlation between predicted and measured strengths (Pearson r = 0.75). This work establishes the first deep learning-enabled RBS design platform for P. denitrificans, offering a robust tool for precise translational regulation and advancing synthetic biology applications in environmental biotechnology.
A data-driven "digital twin" using deep learning-based Molecular Embeddings (MolE) and L1-regularized logistic regression to decode these structural determinants of enzymatic transglycosylation, providing a promising ligand-based virtual screening approach for identifying novel acceptors in enzymatic transglycosylation.
Dong-Ho Seo, Yun-Sang So, S. Yoo· Journal of Agricultural and...· 0 citations
Abstract Motivation Functional characterization of microbiomes often relies on the sequencing of metagenomic DNA extracted from environmental samples, with current approaches using metagenome-assembled genomes (MAGs). Although glycoside hydrolases (GHs) are central to carbon cycling, accurate annotation of GHs in metagenomic datasets remains challenging due to the multidomain architecture of carbohydrate-active enzymes and the prevalence of unassembled short reads due to limitations in the MAG-generation process. Results Here, we present CAZyOGH (CAZymes Open-source GH annotation), a curated reference database for the domain-specific identification of 135 protein domains spanning 99 GH families with well-defined catalytic domain signatures. CAZyOGH focuses on individual GH domains, enabling robust annotation of both assembled and unassembled metagenomic data. We validated CAZyOGH by reanalyzing genomes listed in CAZy db, where predicted GH profiles closely matched reported values. Next, we used CAZyOGH to analyze 12 human gut metagenomes and 12 newly sequenced soil microbiomes to reveal environment-specific GH repertoires. By accurately detecting catalytic domains independent of the genomic context, CAZyOGH improves sensitivity and specificity in short-read metagenomic annotation. This framework provides a scalable and reproducible approach to investigate carbohydrate-active enzymes across ecosystems, advancing our capacity to characterize microbial functional potential in global carbon cycling. Availability and implementation CAZyOGH data is available on figshare (https://figshare.com/projects/CAZyO_GH/267770).
N. Griffin, Alison E Hughes, D. S. Erdody et al.· Bioinformatics Advances· 0 citations
The utility of ML-assisted evolution for engineering Rubisco with improved carboxylation efficiency and potential for enhancing crop productivity is demonstrated.
Julie L. McDonald, Jiachen Lin, Yunlong Zhao et al.· bioRxiv· 0 citations
A targeted mining workflow is developed that screens exclusively plastic-associated datasets through multi-step bioinformatic filtering—integrating catalytic-motif screening, disulfide-topology validation, structural-similarity scoring, and phylogenetic profiling—to recover high-confidence PETase candidates, resulting in a thermostable enzyme that depolymerizes PET across a broad temperature range.
Konstantinos Rigkos, Dimitra S Bezantakou, Kyriakos Antoniadis et al.· bioRxiv· 0 citations
This work demonstrates that integrating cell-free enzyme engineering with machine learning enables opportunities for high-throughput experimental measurements to benchmark and improve protein language models, accelerate design loops, and expand functional exploration within protein families where experimental information is limited.
J. Lazar, Evan Komp, I. Martínez et al.· bioRxiv· 1 citation
DB-IRES serves as a reliable and precise computational tool for predicting IRES elements and enables more in-depth functional investigations of IRES biology and supports wider applications in RNA research and the development of associated therapeutics.