Skip to content
Open access

Large language models enhance annotation of enzymes in metagenomes

Aug 2026 · Science Advances · Vol 12 · 0 citations · 72 references
Medicine

TL;DR

FEDKEA, an enzyme annotation tool leveraging protein language models, and a user-friendly, FEDKEA-based metagenomic pipeline, MEnzMap, which encompasses the entire analysis workflow—from raw data quality control to function prediction and downstream analyses are designed.

Abstract

Metagenomic data have notable biological potential, but their functional interpretation is frequently impeded by incomplete protein function annotations. Accurate enzyme annotation is essential for elucidating the metabolic capabilities of microbial communities within metagenomic datasets. To address this challenge, we developed FEDKEA, an enzyme annotation tool leveraging protein language models, and provided a web platform for its use. In addition, we designed a user-friendly, FEDKEA-based metagenomic pipeline, MEnzMap, which encompasses the entire analysis workflow—from raw data quality control to function prediction and downstream analyses. Applying MEnzMap to human gut metagenomic data from the iHMP2 project, we generated a comprehensive enzyme profile landscape for both healthy individuals and patients with inflammatory bowel diseases. These tools provide an efficient method for the functional annotation of microbial dark matter and facilitate the identification of disease-associated enzymes.

Read PDF

Similar papers

Review Open access Aug 2026

Bioinformatic tools for microbiome analysis: from raw sequences to biological insights

This review presents a practical, workflow-oriented guide to microbiome data analysis, from raw DNA sequence processing to statistical interpretation and biological insight, and highlights emerging technologies, including machine learning methods that are beginning to reshape the field.

Jenna Poelzer, D. Wishart · 0 citations
Open access Aug 2026

Automating scientific annotations for open transcriptomic profiles via multi-stage agents

GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.

Xiaodan Zhang, S. Paithankar, Jing Pu et al. · 0 citations
Open access Jul 2026

WASP: a pipeline for functional annotation prediction based on AlphaFold structural models

WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches, predicting protein functions from AlphaFold structures using network-based structural homology and filling metabolic model gaps by mapping 75-100% of orphan reactions.

Giorgia Del Missier, Kiyan Shabestary, Rodrigo Ledesma-Amaro · 0 citations
Open access Jul 2026

Annotation of glycoside hydrolases in unassembled metagenomes using CAZyOGH

Abstract Motivation Functional characterization of microbiomes often relies on the sequencing of metagenomic DNA extracted from environmental samples, with current approaches using metagenome-assembled genomes (MAGs). Although glycoside hydrolases (GHs) are central to carbon cycling, accurate annotation of GHs in metagenomic datasets remains challenging due to the multidomain architecture of carbohydrate-active enzymes and the prevalence of unassembled short reads due to limitations in the MAG-generation process. Results Here, we present CAZyOGH (CAZymes Open-source GH annotation), a curated reference database for the domain-specific identification of 135 protein domains spanning 99 GH families with well-defined catalytic domain signatures. CAZyOGH focuses on individual GH domains, enabling robust annotation of both assembled and unassembled metagenomic data. We validated CAZyOGH by reanalyzing genomes listed in CAZy db, where predicted GH profiles closely matched reported values. Next, we used CAZyOGH to analyze 12 human gut metagenomes and 12 newly sequenced soil microbiomes to reveal environment-specific GH repertoires. By accurately detecting catalytic domains independent of the genomic context, CAZyOGH improves sensitivity and specificity in short-read metagenomic annotation. This framework provides a scalable and reproducible approach to investigate carbohydrate-active enzymes across ecosystems, advancing our capacity to characterize microbial functional potential in global carbon cycling. Availability and implementation CAZyOGH data is available on figshare (https://figshare.com/projects/CAZyO_GH/267770).

N. Griffin, Alison E Hughes, D. S. Erdody et al. · 0 citations
Open access Jul 2026

Data Independent Acquisition Pipeline for Microbiome Samples (Microbe-DIA)

This work optimized LC–MS/MS acquisition parameters for both DDA and DIA using a model microbiome, demonstrating how DIA enables increased sample throughput without compromising quantitative performance and establishing a scalable and cost-effective pipeline for metaproteomics of complex microbial communities.

Samantha Obermiller, Mary S. Lipton, P. Piehowski et al. · 0 citations