Skip to content
Open access

MEGAHIT k-mer range tuning trades computational efficiency for improved recovery of functional genes across cave sediment and wastewater metagenomes

Jul 2026 · bioRxiv · 0 citations · 30 references
Biology

TL;DR

The study provides a quantitative basis for selecting MEGAHIT k-mer parameters according to whether computational efficiency or functional gene discovery is the primary aim, and indicates that reduced k-mer sets can lower computational cost but may miss biologically relevant functional signal, depending on the dataset and downstream target.

Read PDF

Similar papers

Open access Jul 2026

rbims: an R package for integrative functional profiling and pathway-level discrimination in metagenome-assembled genomes

Metagenomics enables the recovery of metagenome-assembled genomes (MAGs), providing access to the metabolic potential of uncultured microbial communities that drive ecosystem function and biogeochemical cycles. However, as MAGs datasets increase in size and complexity, comparing functional repertoires and identifying ecologically meaningful traits across experimental gradients becomes increasingly difficult. Here, we present rbims, a modular R package for integrative functional profiling of MAGs and metagenomic datasets. rbims supports annotations from KEGG, dbCAN, InterProScan, MEROPS, and PICRUSt2, and enables the calculation of gene presence/absence, raw abundance, and pathway coverage, as well as metadata-informed comparative analyses and publication-ready visualizations. Beyond descriptive profiling, rbims implements an exploratory discriminant framework that combines compositional differential analysis (ALDEx2) with random forest–based feature ranking to prioritize candidate metabolic traits associated with environmental factors. Importantly, it extends gene-level analysis to pathway-level directional bias testing, allowing users to evaluate whether the majority of genes within a metabolic route are consistently enriched toward a given condition. We applied rbims to 42 MAGs recovered from a hydrocarbon enrichment experiment in the North Atlantic Ocean. The workflow identified widespread hexadecane and phenanthrene degradation potential, detected enriched oxidoreductase-related protein families, and revealed a strong pathway-level directional bias toward deep-water MAGs for phenanthrene, naphthalene, and hexadecane degradation pathways. By integrating annotation parsing, quantitative trait analysis, statistical discrimination, and visualization in a reproducible framework, rbims provides a user-friendly platform for functional interpretation in genome-resolved metagenomics.

Karla P. López-Martínez, S. Hereira-Pacheco, Diana Hernández-Oaxaca et al. · 0 citations
Open access Jul 2026

Long-read metagenomics and methylation-based binning support the discovery of antibiotic resistance gene-host associations in complex communities.

By linking ARGs to their wider genetic contexts and hosts, the findings shed light on the previously unrecognized carriers of resistance genes in wastewater, and provides a valuable methodology for early identification of newly arising ARGs and their hosts.

Melina A. Markkanen, Heidi Putkuri, D. Kičiatovas et al. · 0 citations
Open access Aug 2026

Long-read sequencing reveals putatively mobilizable resistance genes and multi-drug resistance plasmids underestimated by short-read metagenomics

While shotgun metagenomics is often used to profile antibiotic resistome in gut microbial communities, few studies have investigated if the choice of sequencing platform and assembly strategy affect what mobile genetic elements and antimicrobial resistance genes are recovered. In this study, we compared three platforms (Illumina, Oxford Nanopore, and PacBio HiFi) and seven assembly strategies on gut metagenomes from cattle, pig, and human as case studies. Long-read assemblies recovered 5- to 7-fold more plasmid sequence than Illumina in cattle and pig (mean 17.0 Mb vs. 3.1 Mb), while Illumina performed comparably in the less diverse human gut where high per-species coverage enabled effective short-read plasmid assembly. Long reads also detected more resistance genes on plasmid contigs. Hybrid assembly results depended on the algorithm: scaffolding-based OPERA-MS preserved long-read contiguity and recovered more plasmid-borne resistance genes, while the short-read-centric metaSPAdes hybrid mode produced fragmented assemblies. After collapsing haplotype redundancy, PacBio HiFi identified 2 and 49 unique multi-drug resistance plasmid lineages in cattle and pig, respectively. On the other hand, only 2 and 4 were identified from Illumina. Long reads also placed far more ARGs in a putative mobilization context (50–73%) compared to 14–21% for short reads. Platform and assembly strategy are thus key variables in mobilome and resistome characterization and should be accounted for in antimicrobial resistance surveillance.

Dabin Jeon, Tatsuya Unno · 0 citations
Review Open access Jul 2026

The MiDAS global genome catalog: 53,501 long-read MAGs representing all core prokaryotic genera in the global activated sludge microbiome

Wastewater treatment relies on complex microbial communities, yet existing genome-resolved references for this essential engineered ecosystem remain dominated by short-read assemblies, limiting genome contiguity and linkage between taxonomic and metabolic function. We applied long-read sequencing to activated sludge from 83 globally distributed plants, reconstructing 53,501 metagenome-assembled genomes to establish the Microbial Database of Activated Sludge (MiDAS) global genome catalog. The catalog encompasses high-quality genomes for 12,047 prokaryotic species, 82% of which are not represented in GTDB release 226, and provides a median of 32 high-quality genomes for each of the 250 core prokaryotic genera previously defined in our MiDAS global 16S rRNA gene survey. This enables analyses of predicted functional traits and their ecological context, for example, we identified two sparsely represented Nitrospiraceae genera with conserved nitrite-oxidation genes that are abundant in higher-temperature wastewater treatment plants. In summary, the MiDAS genome catalog provides a framework for linking taxonomy, metabolism and ecological roles in wastewater treatment systems globally.

Lei Liu, C. Singleton, R. Kirkegaard et al. · 0 citations
Open access Jul 2026

A Sample to Results Workflow for Compositional Analysis of Multiplexed Amplicon Sequencing Experiments

Microbial communities play key roles in the transformation and cycling of elements ranging from required macronutrients to toxic metalloids. Next-generation sequencing has been applied across multiple ecosystems to probe the interplay of microbial community structure and functional potential with respect to elemental cycling. Shotgun metagenomics collects marker gene sequences without amplification and is costly for large numbers of samples and deep coverage. Conversely, amplicon sequencing of taxonomic marker genes, e.g. 16S and 18S rRNA, is cost-effective for large numbers of samples, but provides limited functional insight. A middle ground between the two approaches is needed to analyze community structure and functional potential within a sample while remaining cost-effective with high throughput. To address this need, we developed a standardized workflow for multiplexed amplicon sequencing from sample collection through data analysis for diverse sample types, including freshwater, sediments, and soils, that produces data and publication-ready figures for multiple taxonomic and functional genes for carbon, nitrogen, phosphorus, sulfur, and arsenic cycling for each sample analyzed. The workflow’s utility was shown by analyzing 11 taxonomic and functional gene amplicons sequenced from 25 samples with high technical replicate similarity. The workflow is named CAMASE for Compositional Analysis of Multiplex Amplicon Sequencing Experiments. This proof-of-concept shows that CAMASE economically produces standard amplicon sequencing outputs (ASV/OTU counts and taxonomy, PCA, and relative abundance plots) for hundreds of amplicon by sample combinations and provides specific recommendations for implementation. GRAPHICAL ABSTRACT Samples are collected in a preservative and material collected on filters prior to DNA extraction. Target gene amplicons are produced in parallel with internal barcodes enabling sequencing in a single run followed by compositional data analysis. All wet lab protocols, code markdowns, and templates for required metadata files are available at https://hansonlabgit.dbi.udel.edu/aprange/CAMASE. Created in BioRender. Bennett, A. (2026) https://BioRender.com/ymnojt0

Alexa J. Bennett, Ryan M. Moore, Craig W. Herbold et al. · 0 citations
Open access Jul 2026

metaSMASH: Scalable Biosynthetic Gene Cluster Detection for Large Metagenomic Assemblies

antiSMASH is widely used for biosynthetic gene cluster (BGC) detection and annotation, but its standard workflow is poorly suited to large metagenomic assemblies, where massive contig counts create severe runtime bottlenecks and complicate downstream result exploration. We present metaSMASH, a re-engineered fork of antiSMASH for metagenome-scale BGC analysis. metaSMASH preserves the original antiSMASH detection and annotation logic while introducing streaming, memory-bounded execution, record-level parallelisation, optional output filtering, and an interactive dashboard for large result sets. Across 25 benchmark metagenome datasets, metaSMASH reproduced identical BGC detection results while dramatically reducing computational cost. Relative to the default antiSMASH configuration, metaSMASH was a geometric-mean 38× faster. It also outperformed an ad hoc chunked antiSMASH workflow: in the default configuration it achieved a geometric-mean 2.9 × speed-up and 1.7 × lower peak memory, and with extended-analysis modules enabled it was 2.7 × faster and used 3.1 × less memory while completing all datasets, whereas the ad hoc workflow ran out of memory on the two largest assemblies. By substantially reducing the computational burden of large-scale metagenome analysis without sacrificing result equivalence, metaSMASH makes routine mining of assembled metagenomes more practical and provides a scalable foundation for natural product discovery from complex microbial communities. Graphical Abstract

Caner Bağcı, K. Blin, N. Ziemert · 0 citations