GeneHunt2 is presented, a scalable framework for multidomain annotation of CAZymes that integrates curated HMM profiles from dbCAN and Pfam into a unified, deduplicated database, enabling systematic identification of both CAZy and non-CAZy domains.
Abstract
Carbohydrate-active enzymes (CAZymes) are central to carbohydrate metabolism, yet their functional annotation is typically restricted to catalytic CAZy domains, overlooking the broader multidomain architectures in which these domains operate. Here, I present GeneHunt2, a scalable framework for multidomain annotation of CAZymes that integrates curated HMM profiles from dbCAN and Pfam into a unified, deduplicated database, enabling systematic identification of both CAZy and non-CAZy domains. After confirming the robust recovery of CAZy domain assignments using GeneHunt2, I investigated the detailed multidomain architecture of over 3.75 million CAZyme sequences: more than 40% were multidomain, and non-CAZy partner domains constituted a substantial fraction of detected partner domains. Using a quantitative framework that combines co-occurrence enrichment, domain adjacency, positional bias, and partner-specificity scoring, I next distinguished family-specific modules, auxiliary domains, and promiscuous partners. This approach recapitulates known CAZyme-CBM relationships and extends beyond CAZy definitions by identifying numerous Pfam domains including many domains of unknown function (DUFs) that are specifically and non-randomly associated with particular CAZy families. By enabling reproducible, multidomain-aware annotation, GeneHunt2 facilitates data-driven hypotheses about poorly characterized domains and widens the functional interpretation of carbohydrate-active proteins beyond their catalytic cores.
Abstract Motivation Functional characterization of microbiomes often relies on the sequencing of metagenomic DNA extracted from environmental samples, with current approaches using metagenome-assembled genomes (MAGs). Although glycoside hydrolases (GHs) are central to carbon cycling, accurate annotation of GHs in metagenomic datasets remains challenging due to the multidomain architecture of carbohydrate-active enzymes and the prevalence of unassembled short reads due to limitations in the MAG-generation process. Results Here, we present CAZyOGH (CAZymes Open-source GH annotation), a curated reference database for the domain-specific identification of 135 protein domains spanning 99 GH families with well-defined catalytic domain signatures. CAZyOGH focuses on individual GH domains, enabling robust annotation of both assembled and unassembled metagenomic data. We validated CAZyOGH by reanalyzing genomes listed in CAZy db, where predicted GH profiles closely matched reported values. Next, we used CAZyOGH to analyze 12 human gut metagenomes and 12 newly sequenced soil microbiomes to reveal environment-specific GH repertoires. By accurately detecting catalytic domains independent of the genomic context, CAZyOGH improves sensitivity and specificity in short-read metagenomic annotation. This framework provides a scalable and reproducible approach to investigate carbohydrate-active enzymes across ecosystems, advancing our capacity to characterize microbial functional potential in global carbon cycling. Availability and implementation CAZyOGH data is available on figshare (https://figshare.com/projects/CAZyO_GH/267770).
N. Griffin, Alison E Hughes, D. S. Erdody et al.· Bioinformatics Advances· 0 citations
Gcn5-related N-acetyltransferases (GNATs) are considered a "megafamily" that ranks among the most structurally and sequentially diverse superfamilies in the CATH (Class, Architecture, Topology, Homology) database. In this vast superfamily, several types of protein functions have been explored throughout evolution, yet the evolutionary pathways that led to such diversity remain poorly understood. To investigate these concepts further, we selected a functionally distinct GNAT subgroup called polyamine N-acetyltransferases (PAATs). These enzymes acetylate polyamines that are crucial for cellular homeostasis. While PAATs from different domains of life catalyze the same reaction, their residue conservation patterns, oligomeric states, and presence of allosteric sites vary. Despite their biological importance, many putative PAATs remain uncharacterized, limiting our ability to infer evolutionary relationships, understand how functional properties emerged, and appreciate the extent of their structural diversity and substrate specificity. Here, we present a characterization of a large subset of PAAT enzymes, including their likely oligomeric states, functional site properties, and experimental functions.
J. Roca-Martínez, Hazel N. Leiva Martel, Jialin Yin et al.· Structure· 0 citations
ABSTRACT The rapid expansion of bacterial genome databases presents significant opportunities for functional discovery, as a large fraction of genes and protein domains remain uncharacterized. Analyzing genomic context and domain architecture is a powerful approach for functional inference, but existing tools often lack the scalability and integrated workflow required for high-throughput analysis. To address this, we developed Pandoomain, a Snakemake pipeline that automates the acquisition of genomes from the National Center for Biotechnology Information, identifies proteins of interest using hidden Markov models (HMMs), and performs systematic domain annotation and gene neighborhood analysis. We demonstrate the utility of Pandoomain through a comprehensive analysis of the poorly characterized pre-toxin TG (PT-TG) domain across 347,289 bacterial genomes. Our analysis revealed 10,226 PT-TG-containing proteins organized into 312 unique domain architectures, highlighting their association with diverse interbacterial antagonistic systems, including the Type VI secretion, Type VII secretion, and contact-dependent inhibition systems. By leveraging genomic context, we identified a novel variant of the WXG trafficking domain, termed W10XG, and subsequently discovered 24 new families of associated toxin domains. We experimentally validated six of these toxins, confirming that all six are neutralized by their cognate immunity proteins. Pandoomain is an accessible tool that enables systematic, large-scale exploration of protein domains, and our analysis of the PT-TG domain provides a rich resource for future investigations into the mechanisms and evolution of bacterial antagonism. IMPORTANCE The rapid growth of bacterial genomic data presents a major hurdle for scientists seeking to understand the functions of newly discovered genes and proteins. To address this issue, we created Pandoomain, a powerful, accessible software tool that automates large-scale analysis of genetic information across hundreds of thousands of genomes. Using Pandoomain, we investigated a poorly understood family of proteins involved in bacterial competition, revealing novel protein domain architectural diversity. This led to the discovery of 24 new families of toxins predicted to be used by bacteria to attack their competitors, and we experimentally confirmed the toxic activity of six of them. Our work provides the scientific community with a robust tool to accelerate functional discovery and offers new insights into the evolution of bacterial conflicts, which may provide insights into the compositional dynamics of microbial communities and support methods to engineer their composition. The rapid growth of bacterial genomic data presents a major hurdle for scientists seeking to understand the functions of newly discovered genes and proteins. To address this issue, we created Pandoomain, a powerful, accessible software tool that automates large-scale analysis of genetic information across hundreds of thousands of genomes. Using Pandoomain, we investigated a poorly understood family of proteins involved in bacterial competition, revealing novel protein domain architectural diversity. This led to the discovery of 24 new families of toxins predicted to be used by bacteria to attack their competitors, and we experimentally confirmed the toxic activity of six of them. Our work provides the scientific community with a robust tool to accelerate functional discovery and offers new insights into the evolution of bacterial conflicts, which may provide insights into the compositional dynamics of microbial communities and support methods to engineer their composition.
E. Soto, Adam Oliver, Marcos H. de Moraes· mSystems· 0 citations
The diazo group is a highly valuable functional group in organic synthesis; however, its conventional preparation requires harsh conditions. Here, we provide a systematic analysis of ATP-dependent diazotases from actinomycetes that catalyze the condensation of nitrite with aromatic amines. Sequence similarity network analysis of AMP-binding enzymes revealed two families (type Ia/Ib), with type Ib further classified into 14 groups. In vitro assays confirmed the diazotization activities of the diazotases from 7 groups in the type Ib family, each displaying distinct substrate preferences. Remarkably, diazotases in group 3 exhibited broad substrate specificity and superior catalytic efficiency. The dimeric structure of a representative enzyme, Mco01_40450, was determined at 3.08 Å resolution by cryogenic electron microscopy with C2 symmetry. Subsequent symmetry expansion and 3D classification revealed the structure with AMP, pyrophosphate, and 4-aminohydrocinnamic acid bound to the active site at a resolution of 3.16 Å. Expansion of the substrate-binding pocket and its entrance via amino acid substitution enabled the reaction of bulkier aromatic amines. Unlike classical chemical diazotization, Mco01_40450 operates under mild aqueous conditions and is compatible with sensitive protecting groups. Our findings highlight the structural and functional diversity of diazotases and establish group 3 diazotases as promising, engineerable biocatalysts for selective and efficient diazo installation.
Jiayu Ning, Seiji Kawai, Yohei Katsuyama et al.· Journal of the American Chem...· 0 citations
BACKGROUND
The discovery of novel biocatalysts for the sustainable valorization of complex biomass feedstocks remains a significant challenge. Domain-centric exploration of characterized CAZyme families offers a promising but underexplored strategy for identifying enzymes with unusual architectures and potentially expanded substrate specificities.
RESULTS
Systematic analysis of archaeal glycoside hydrolase family 18 (GH18) chitinases using the CANDy domain annotation pipeline led to the identification of TcChi from Thermococcus chitonophagus, a multidomain enzyme combining a GH12 and a GH18 catalytic domain alongside two carbohydrate-binding modules. Given that T. chitonophagus also encodes dedicated standalone cellulases and chitinases, we hypothesized that this multidomain assembly may have evolved a broader functional range than either composing domain alone. Biochemical assays of truncated constructs confirmed this hypothesis: the GH18 domain hydrolyzed chitin, chitosan, and β-1,3-glucan, marking the first report of β-1,3-glucanase activity (EC 3.2.1.58) in a GH18 chitinase, while the GH12 domain exhibited strong cellulase activity alongside unexpected chitosanase activity (EC 3.2.1.132), extending the known functional range of this family. Both domains demonstrated high thermostability consistent with the hyperthermophilic origin of T. chitonophagus.
CONCLUSIONS
TcChi is a thermostable, multifunctional biocatalyst capable of degrading chitin, chitosan, cellulose, and β-1,3-glucan from a single protein scaffold, making it a promising candidate for consolidated biomass deconstruction and waste valorization. These findings also demonstrate that domain-centric analysis of CAZyme families is an effective strategy for uncovering hidden functional diversity in well-characterized enzyme families.
Alex Windels, S. Dhaene, Tom Desmet· Biotechnology for Biofuels a...· 0 citations
In this study, we developed and applied an integrative computational workflow for the systematic identification and prioritization of candidate allosteric pockets across all four glycolytic enzymes: three from Staphylococcus aureus-phosphoglucose isomerase (PGI), phosphoglycerate kinase (PGK), and enolase- and one representative hexokinase from Plasmodium vivax, included due to the absence of an experimentally determined three-dimensional structure for the S. aureus ortholog. Solvent mapping using FTMap and FTMove across oligomeric ensembles revealed multiple high-confidence cavities predominantly located at subunit interfaces, in addition to canonical catalytic sites. Independent evaluation with CavityPlus supported the presence and druggability of these pockets. Hexokinase presented 12 interface-associated pockets that emerged only upon oligomer formation and remained stable across 300 FTMove-derived conformers. PGI and PGK displayed interface- and hinge-associated cavities linked to known global motions, while the octameric enolase showed prominent central and peripheral inter-dimer pockets. Candidate pockets were subsequently evaluated using CorrSite, ESSA, PASSer, and AlloSigMA to assess features associated with allosteric communication, energetic coupling, and protein dynamics. High-confidence candidate sites were prioritized based on consensus across these complementary computational approaches. Across all four enzymes, interface-localized pockets consistently emerged as promising candidate regulatory regions, suggesting that protein-protein interfaces may represent valuable targets for allosteric modulation. Several predicted pockets, particularly in PGI and enolase, exhibited low sequence and structural similarity to the corresponding human homologs, indicating their potential for selective inhibitor design. Overall, this integrated computational framework provides a systematic strategy for identifying and prioritizing candidate allosteric pockets for future structural, biochemical, and structure-based drug discovery.
Defne Alnıgeniş, Florihana Brina, Ilknur Kocal et al.· Biochimica et Biophysica Act...· 1 citation