Metagenomics enables the recovery of metagenome-assembled genomes (MAGs), providing access to the metabolic potential of uncultured microbial communities that drive ecosystem function and biogeochemical cycles. However, as MAGs datasets increase in size and complexity, comparing functional repertoires and identifying ecologically meaningful traits across experimental gradients becomes increasingly difficult. Here, we present rbims, a modular R package for integrative functional profiling of MAGs and metagenomic datasets. rbims supports annotations from KEGG, dbCAN, InterProScan, MEROPS, and PICRUSt2, and enables the calculation of gene presence/absence, raw abundance, and pathway coverage, as well as metadata-informed comparative analyses and publication-ready visualizations. Beyond descriptive profiling, rbims implements an exploratory discriminant framework that combines compositional differential analysis (ALDEx2) with random forest–based feature ranking to prioritize candidate metabolic traits associated with environmental factors. Importantly, it extends gene-level analysis to pathway-level directional bias testing, allowing users to evaluate whether the majority of genes within a metabolic route are consistently enriched toward a given condition. We applied rbims to 42 MAGs recovered from a hydrocarbon enrichment experiment in the North Atlantic Ocean. The workflow identified widespread hexadecane and phenanthrene degradation potential, detected enriched oxidoreductase-related protein families, and revealed a strong pathway-level directional bias toward deep-water MAGs for phenanthrene, naphthalene, and hexadecane degradation pathways. By integrating annotation parsing, quantitative trait analysis, statistical discrimination, and visualization in a reproducible framework, rbims provides a user-friendly platform for functional interpretation in genome-resolved metagenomics.
This review presents a practical, workflow-oriented guide to microbiome data analysis, from raw DNA sequence processing to statistical interpretation and biological insight, and highlights emerging technologies, including machine learning methods that are beginning to reshape the field.
Jenna Poelzer, D. Wishart· Frontiers in Microbiology· 0 citations
Phylogenomic inference requires empirical datasets that capture the diversity of molecular evolutionary processes, including heterogeneity in substitution processes and phylogenetic signal, yet existing resources rarely provide standardized, per-locus alignment, model, and tree metrics. We present MsaTM-DB, a curated database comprising 965,545 loci from 420 eukaryotic phylogenomic studies, each annotated with 35 features spanning alignment properties, substitution parameters, and gene tree metrics. These data enable systematic investigation of heterogeneity in phylogenetic signals across diverse evolutionary contexts. An integrated pipeline and interactive R Shiny platform support distributional analyses, correlation exploration, and empirically informed simulations. Using this database, we illustrate its downstream potential through exploratory analyses showing that alignment- and tree-derived features can help evaluate factors influencing phylogenetic support. MsaTM-DB provides an extensive empirical foundation for benchmarking phylogenetic methods, guiding marker selection, and developing data-driven evolutionary models. Our database is available online at the GitHub repository (https://github.com/xtmtd/MSA-and-tree-metrics-exploration).
The increasing availability of genomic and metagenomic data has created significant opportunities to explore microbial diversity, biosynthetic potential and functional traits. However, comprehensive and comparative genome analysis often requires integrating multiple independent tools, making large-scale studies challenging to implement and manage. Here, we present Bacterial Genome eXplorer (BGX), an integrated and scalable pipeline that streamlines large-scale genome analysis and exploration of biosynthetic potential. BGX integrates different analytical tools into six major stages: (i) genome retrieval and assembly, (ii) genome quality assessment, (iii) antimicrobial resistance (AMR) gene profiling, (iv) annotating Biosynthetic gene clusters (BGCs) and novelty assessment, (v) bioactivity predictions, and (vi) clustering and networking analysis. In addition, BGX provides a user-friendly interactive interface to facilitate data exploration and interpretation. We demonstrated the versatility and scalability of BGX through large-scale analysis of two independent datasets: 248 genomes from the One Day One Genome (ODOG) initiative and 153 publicly available genomes from NCBI. This analysis enabled the comprehensive characterisation of genome quality, AMR determinants, biosynthetic potential and candidate bioactive metabolites across two datasets. BGX is distributed as a Docker container that simplifies installation, enables reproducible data processing, and supports pipeline execution across different computational environments. The modular and reproducible architecture of the BGX pipeline provides an effective framework for large-scale genome mining, genomic surveillance, and accelerates the discovery and prioritisation of novel secondary metabolites. BGX is now accessible at https://bgx.nabi.res.in
Abstract Metagenomic sequencing is transforming diverse areas of health and biological sciences, including pathogen surveillance, clinical diagnostics, and microbiome research. However, the inherent complexity of metagenomic data limits most computational tools to species-level classification and abundance estimation, overlooking within-species genetic diversity that drives key phenotypes. We present metaWEPP, a novel computational pipeline that achieves near-haplotype resolution in metagenomic analysis for species with adequate representation in reference genome biobanks and having sufficient sequencing depth and genome coverage. Specifically, metaWEPP assigns sequencing reads to species using standard taxonomic classifiers, phylogenetically places them onto species-specific mutation-annotated trees of publicly available sequences, and selects the haplotypes that best explain the sample. It also reports unaccounted alleles indicative of novel variants and provides an interactive dashboard for read-level visualization. Applied to diverse metagenomic and mixed-genome samples from prior studies, metaWEPP produced concordant species-level results, while revealing finer lineage- and haplotype-level insights not captured by existing tools. On various clinical samples, metaWEPP identified infecting pathogens and additionally provided credible lineage- and haplotype-level information that can support clinical decision-making. On wastewater samples, metaWEPP uncovered previously undetected haplotype clusters of epidemiological relevance. These findings demonstrate metaWEPP’s ability to advance various clinical, epidemiological, and research applications with deeper, actionable insights.
Pranav Gangwar, Qiwen Xu, Jaden Seangmany et al.· NAR Genomics and Bioinformat...· 0 citations
Application to a human sample from a patient with type 2 Diabetes Mellitus recovered a dysbiotic signature consistent with the literature, including reduced Firmicutes abundance, elevated Bacteroidetes and Proteobacteria, and a predominance of clinical associations within metabolic and gastrointestinal categories.
Rodrigo Lima Andrade, Tayná da Silva Fiúza, J. Kroll et al.· bioRxiv· 0 citations
Metabolic interactions govern gut microbiome assembly, yet their functional rules remain obscured by genomic incompleteness and fragmentation. Here, we leverage 1,150 complete genomes to construct genome-scale metabolic models, demonstrating that draft assemblies introduce systematic artifacts and omit critical transport functions. We observe that genomic traits and niche specialization, rather than random association, shape microbial metabolic competition and complementarity. Interaction asymmetry stratifies strains into four ecological groups, including active players, resource predators, resource utilizers, and resource contributors, with distinct signatures of metabolite exchange, competition, and secondary metabolism. In inflammatory bowel disease, these groups show subtype-specific temporal instability, and group-specific dysbiosis predicts clinical phenotypes better than the whole-community profiles. Keystone features derived from integrated metabolic interaction and co-occurrence networks also improve cross-validated disease classification. Together, these findings connect genome completeness with microbial ecological organization and provide a framework for linking metabolic interactions to microbiome-associated disease.
Yu-He Gu, Haoyu Wang, Jin-Long Yang et al.· Cell Reports· 0 citations