Skip to content
Open access

MGM2 as a Unified Foundation Model for Microbiome World Exploration

Jul 2026 · bioRxiv · 0 citations · 39 references
Biology

TL;DR

MGM2 provides a sequence-aware and interpretable representation layer for microbiome classification, paired-community prediction, forecasting and feature discovery and a 4,096-feature sparse autoencoder atlas resolves taxonomic, abundance, ecological and technical signals.

Abstract

Microbiomes are information-rich biological systems, yet most computational analyses still reduce communities to cohort-specific abundance tables. Here we introduce MGM2, a multimodal foundation model pretrained on 1,821,291 MicrobeAtlas samples and 225,067 OTUs clustered at 99% sequence similarity. MGM2 couples NTv3-derived microbial sequence embeddings with abundance conditioning and community-semantic alignment to learn transferable sample– and token-level representations. Frozen MGM2 representations outperformed DeepPhylo by 0.06–0.21 macro-AUROC across five temporally held-out MGnify hierarchy levels, with the largest gains for rare and fine-grained labels. In fecal microbiota transplantation, MGM2-XLarge achieved a response ROC AUC of 0.79 and reduced post-transplant Bray-Curtis distance by 15% relative to the recipient baseline. The same representation supported ASV-level trend forecasting across 24 wastewater treatment plants. Sparse autoencoder analysis resolved MGM2-XLarge token states into a 4,096-feature dictionary spanning taxonomic identity, abundance state, ecological context and technical variation. MGM2 therefore provides a sequence-aware and interpretable representation layer for microbiome classification, paired-community prediction, forecasting and feature discovery. Highlights ● MGM2 integrates sequence, abundance and community semantics through pretraining on 1.82 million microbiome samples. ● Frozen MGM2 improved macro-AUROC over DeepPhylo by 0.06–0.21 across five temporally held-out MGnify levels. ● MGM2-XLarge reached a response ROC AUC of 0.79 and reduced post-FMT Bray-Curtis distance by 15%. ● A 4,096-feature sparse autoencoder atlas resolves taxonomic, abundance, ecological and technical signals.

Read PDF

Similar papers

Open access Aug 2026

GTX-GUT: A Standardized Metagenomic Workflow for Gut Microbiome Profiling and Clinical Associations

Application to a human sample from a patient with type 2 Diabetes Mellitus recovered a dysbiotic signature consistent with the literature, including reduced Firmicutes abundance, elevated Bacteroidetes and Proteobacteria, and a predominance of clinical associations within metabolic and gastrointestinal categories.

Rodrigo Lima Andrade, Tayná da Silva Fiúza, J. Kroll et al. · 0 citations
Open access Jul 2026

Predicting the seed microbiome using phylogeny-driven machine learning

The composition of the seed-associated bacterial microbiome can reflect host evolutionary relationships, a pattern consistent with phylosymbiosis. While machine learning offers new opportunities to predict microbial community composition, existing models often require prior microbial profiles or environmental variables, limiting their application to unsampled hosts. Here, we tested whether plant nuclear internal transcribed spacer (ITS) sequences, used as a marker of host relatedness, can predict species-level seed-associated bacterial communities using 16S rRNA data from 61 plant species. We introduced customized machine learning models that use sequence-based Hamming distances to capture plant host relatedness. Among the tested models, the Hamming Distance-based k-Nearest Neighbor model (HD-KNN) achieved the highest overall predictive accuracy, yielding an average Jensen-Shannon divergence (JSD) of 0.276 between observed and predicted microbiome profiles. HD-KNN performed particularly well within densely sampled host groups, including Brassicaceae and Poaceae, where closely related reference species were available. In contrast, Hamming Distance-based Gaussian Process Regression (HD-GPR) showed slightly better performance for phylogenetically isolated species, suggesting that model performance depends on host representation within the training dataset. Our framework demonstrates that plant nuclear ITS-derived host relatedness carries a partial predictive signal for seed-associated bacterial microbiome composition. These results provide a foundation for low-input predictive modelling of seed-associated bacteria and may help prioritise microbiome predictions for unsampled plant species when closely related reference species are available. However, our conclusions are strictly limited to seed-associated bacterial communities and should not be directly generalized to fungal communities or other plant compartments, such as the rhizosphere or phyllosphere, which may be shaped by different environmental filtering mechanisms.

Julia Herbinger, D. Ramakrishnan, Jannik Reißfelder et al. · 0 citations
Open access Aug 2026

Joint-RPCA: domain-aware multi-omics integration for systems microbiology.

Integrating multi-omics data is essential for microbiome research, as microbial communities are shaped by and respond to interdependent processes, including taxonomic composition, metabolite production and utilization, and gene expression. However, accurately capturing ecosystem-wide patterns across these modalities is statistically challenging due to differences in scale, sparsity, and compositionality. While a growing number of multi-omics methods have emerged, they differ in their mathematical objectives and modeling assumptions, which in turn shape how biological patterns are represented and interpreted. This underscores the need for tools that explicitly account for the statistical properties of microbial ecosystems. Here, we present Joint Robust Principal Component Analysis (Joint-RPCA), a method designed with these statistical properties in mind and broadly applicable to multi-omics settings with similar challenges. Built on the OptSpace matrix completion framework, Joint-RPCA assumes an underlying shared low-rank structured component across modalities to identify shared variation and cross-modal associations from matched samples. Within this setting and under these statistical assumptions, Joint-RPCA showed stronger performance than the benchmarked general-purpose methods in phenotype separation and feature association tasks, achieving up to sixfold improvement in classification accuracy and over 100-fold faster runtimes. Applied to real-world datasets, including the Integrative Human Microbiome Project (iHMP), mammalian gut microbiomes, and decomposition studies, Joint-RPCA reveals replicable and interpretable multi-omic patterns, offering a scalable and domain-aware solution for systems-level microbiome analysis. Joint-RPCA is available in both Python ( https://github.com/biocore/gemelli ) and R ( https://bioconductor.org/packages/mia ).

Bianca Cordazzo Vargas, C. Martino, A. Dilmore et al. · 1 citation
Open access Aug 2026

Ecological Network Inference Reveals 737 Cross-Kingdom Associations Structuring Human Microbiomes

The human microbiome is a complex, multikingdom ecosystem where bacteria and fungi cohabit and interact. Despite their ecological and clinical significance, cross-kingdom dynamics remain poorly characterized due to dominant single-kingdom research approaches. To understand the principles structuring multi-kingdom microbial communities, we applied the sparse inference method SpiecEasi to 45 publicly available samples from the gastrointestinal tract, skin, and oral cavity. Bacterial (16S rRNA) and fungal (ITS) sequencing data were processed using QIIME2, managed in phyloseq, and co-occurrence networks were inferred via SpiecEasi with Meinshausen– Bühlmann estimation. To validate robustness, we employed SparCC as a secondary inference method and performed 100 bootstrap iterations. Body site stratification controlled for environmental confounders. Our analysis revealed a microbial network of 5,023 taxa (5,020 bacterial, 3 fungal) connected by 30,478 significant associations. Crucially, we identified 737 robust bacterial–fungal interkingdom interactions (689 positive, 48 negative) confirmed by both inference methods. The network exhibited sparse connectivity (density = 0.0024) and modular structure (modularity = 0.45). Hub analysis identified 15 keystone taxa, including Bacteroides uniformis and Faecalibacterium prausnitzii. Interaction patterns were body-site-specific (P < 0.001), with the gastrointestinal tract showing the highest interkingdom connectivity (385 edges). This study provides systematic evidence that bacterial–fungal interactions are abundant and integral to human microbiome architecture. The discovery of 737 cross-kingdom associations challenges the prevailing single-kingdom paradigm and advocates for an integrated multikingdom perspective. These interactions, particularly those mediated by keystone hubs, represent novel targets for microbiome-based therapeutics and diagnostics. Importance This study challenges the prevailing single-kingdom paradigm in microbiome research by demonstrating that bacterial–fungal interactions are abundant and integral to human microbiome architecture. The discovery of 737 cross-kingdom associations across three body sites provides a foundational resource for understanding multikingdom microbial ecology. The identification of keystone bacterial hubs—particularly Bacteroides uniformis and Faecalibacterium prausnitzii—as central connectors in interkingdom networks opens new avenues for microbiome-based therapeutics and diagnostics. Our integrated analytical framework, combining SpiecEasi and SparCC with body site stratification, offers a robust methodological template for future cross-kingdom studies.

A. Babaei, Seyed davar Siadat · 0 citations
Open access Aug 2026

The multi-omics fallacy in microbiome science

Artificial intelligence and machine-learning-assisted multi-omics have expanded the scale and ambition of microbiome research, but they have also sharpened an older interpretive problem. Biologically plausible structure is too easily mistaken for biological explanation. This Perspective defines the multi-omics fallacy as claim inflation that occurs when integrating microbial, host, environmental, spatial, and clinical data is assumed to move interpretation from association toward verified mechanism without a corresponding gain in measurement, localization, temporal resolution, functional linkage, or perturbation. The risk is not that computational integration lacks value. It is that predictions, imputations, inferred pathways, feature attributions, and cross-layer networks can acquire mechanistic authority before their biological status has been established. In microbiome science, where stool readouts, taxonomic abundance, inferred function, and predicted metabolites often serve as proxies for host-microbial interaction, this slippage can make uncertain claims appear more complete than the evidence allows. Computational confidence can amplify structured artifact when systematic error becomes learnable. This Perspective proposes a model-to-mechanism burden of proof that distinguishes prediction from explanation, imputation from observation, attribution from causality, cross-layer coherence from mechanism, and diagnostic performance from biological validity. This framework is intended to strengthen, not constrain, computational microbiome science by clarifying which outputs support classification or hypothesis generation and which require direct measurement, localization, temporal analysis, functional validation, or perturbation. Used this way, computational models can help expose uncertainty, prioritize experiments, identify fragile claims, and sharpen biological questions. The result would be a more powerful form of computational microbiome science, one in which models do not stand in for mechanisms but guide the work needed to earn them.

Rebecca Lewandowski · 0 citations
Open access Aug 2026

rCCLasso: a robust framework for microbial correlation network analysis reveals age-related microbial dynamics

The human gut microbiome continues to evolve beyond early adulthood, yet most microbiome aging studies focus on changes in individual taxa or overall diversity, leaving microbial interaction dynamics largely unexplored. Correlation-based microbial networks offer an interpretable framework for studying such interactions but are challenging to estimate from compositional microbiome data. Although compositionality-aware methods such as CCLasso provide principled multivariate inference, we identify a previously overlooked limitation: sensitivity to random seeds, which leads to unstable correlation estimates and irreproducible significance assessments. To address this issue, we propose Robust CCLasso (rCCLasso), a statistically rigorous framework that stabilizes microbial correlation estimation by integrating CCLasso outputs across multiple runs. rCCLasso aggregates sparse correlation estimates using median-based integration with positive-definite projection and combines run-specific inference through the Cauchy combination test with an additional stability criterion to control type-I error. The method is naturally parallelizable and computationally scalable. Simulation studies demonstrate that rCCLasso improves inferential stability, type-I error control, and power relative to the original CCLasso. Applying rCCLasso to data from over 4,000 healthy adults in the American Gut Project (ages 18--101), we uncover age-related microbial network dynamics, characterized by marked fluctuations from early to mid-adulthood, followed by a relatively stable phase and a substantial decline in network strength in the elderly group. Together, these results establish rCCLasso as a robust and interpretable framework for studying microbial networks in aging research.

Tianyi Xie, Jie Zhou, Yue Wang · 0 citations