Skip to content
Open access

A generalized supervised contrastive learning framework for integrative multi-omics prediction models

Aug 2026 · Frontiers in microbiomes · Vol 5 · 0 citations · 34 references
Medicine

TL;DR

MB-SupCon-cont improves prediction accuracy by incorporating a generalized contrastive loss function that defines similarity and dissimilarity for continuous responses using three distance-based weighting methods, and provides superior representation learning and improves data visualization in lower-dimensional spaces.

Abstract

Advancements in multi-omics research have demonstrated the potential of integrating human microbiome and metabolomics data to better understand physiological processes and improve prediction accuracy in studies of human health. While conventional models utilizing single-omics data provide valuable perspectives, they often fail to capture the complexity of biological systems. Recent developments in supervised contrastive learning frameworks have enhanced predictive performance for categorical responses, yet limitations persist in extending these methods to continuous outcomes. A robust model capable of addressing these gaps could significantly enhance multi-omics predictions and provide new insights into complex biological interactions. Here, we present MB-SupCon-cont, a novel supervised contrastive learning framework designed for both categorical and continuous responses in multi-omics data. MB-SupCon-cont improves prediction accuracy by incorporating a generalized contrastive loss function that defines similarity and dissimilarity for continuous responses using three distance-based weighting methods. Through simulation studies and two real-world datasets for Type 2 Diabetes (T2D) and High-Fat Diet (HFD), we demonstrate that MB-SupCon-cont consistently achieves lower prediction errors than tuned conventional models, canonical correlation analysis, and autoencoder baselines, with most reaching statistical significance. We further provide a validation-based rule for selecting the weighting method and show that the learned embeddings align more closely with the response and recover known microbe and metabolite associations. The framework also provides superior representation learning and improves data visualization in lower-dimensional spaces. These findings suggest that MB-SupCon-cont is a powerful tool for general multi-omics prediction and may have broad applicability in biomedical research.

Read PDF

Similar papers

Open access Aug 2026

When Clinical and Metabolomics Data Work Together: A Comparative Framework for Multimodal Disease Classification across Machine and Deep Learning

Disease classification using clinical and metabolomics data increasingly relies on multimodal integration, yet the complementary and comparative contributions of these modalities remain poorly understood. Most existing frameworks prioritize predictive performance without systematically examining how data modalities and modeling paradigms influence classification outcomes. Consequently, the relative value of individual modalities versus their integration, particularly in terms of model behavior, robustness, and interpretability, remains poorly characterized across machine learning (ML) and deep learning (DL) approaches. Here, we present a modality-aware comparative framework that enables direct, side-by-side evaluation of clinical-only, metabolomics-only, and combined modeling strategies within a unified pipeline. Unlike existing tools designed primarily for multiomics integration, this framework explicitly assesses when and how each modality contributes value across diverse classification scenarios. It supports the efficient use of available data while enabling a systematic comparison of model performance, stability, and interpretability across ML and DL methods. Rather than introducing a new classifier model, this work delivers a unified benchmarking workflow for evidence-based decision-making of modeling strategies in small-sample clinical metabolomics settings. Applied to two glomerulonephritis cohorts representing clinically-driven and metabolomics-driven classification settings, the framework revealed that the dominant contributing modality shifted by scenario, reflecting differences in disease context. These scenarios reflect realistic situations where discriminative signals arise from clinical variables, metabolic alterations, or their combination, as well as more complex cases involving subclass discrimination with overlapping profiles. While several models achieved comparable predictive accuracy, they differed in feature ranking stability, sampling sensitivity, and tendency to overfitting. Overall, this framework facilitates transparent, evidence-based selection of modeling strategies and data modalities suited to data complexity, sample size, and research objectives. Source code is available at https://github.com/kwanjeeraw/MMFramework.

K. Wanichthanarak, Kassaporn Duangkumpha, Nichapa Kleebkomut et al. · 0 citations
Review Open access Aug 2026

Multi‐omics–driven precision medicine

Abstract Precision medicine is increasingly constrained not by a lack of molecular data but by the absence of frameworks that can translate multidimensional biological information into actionable clinical decisions. Multi‐omics‐driven precision medicine (MODPM) addresses this lack by integrating genomics, epigenomics, transcriptomics, proteomics, metabolomics, microbiome, and clinical context into a multiscale framework that links molecular mechanisms, tissue organization, and patient trajectories. In this review, we propose a conceptual framework for MODPM and examine how advances in multi‐omics technologies, artificial intelligence (AI), and foundation models are reshaping disease modeling, drug development, and precision intervention. We summarize the biological contributions of major omics layers and discuss how AI supports cross‐modal representation learning, contextual modeling, and perturbation‐aware prediction. We highlight drug development as a key translational application of MODPM and further discuss its clinical relevance across three major disease contexts: cancer, autoimmune diseases, and metabolic disorders, including cardiometabolic and renal–metabolic diseases. These examples illustrate how MODPM can support target discovery, disease endotyping, treatment response prediction, and clinical monitoring by analyzing shared mechanisms such as immune dysregulation, metabolic remodeling, chronic inflammation, tissue microenvironmental changes, and gene–environment interactions. Across these settings, MODPM enables finer molecular stratification, the identification of pathway‐dominant disease states, improved response prediction, and dynamic treatment monitoring. We also discuss key barriers to implementation, including data heterogeneity, limited cohort diversity, polygenic complexity, workflow constraints, cost, and ethical issues related to privacy, consent, and data ownership. Overall, the value of MODPM lies not in stacking additional data layers but in building a multiscale, continuously learnable framework to link biological heterogeneity to clinically interpretable and actionable decisions.

Huibo Li, Zhe Zhao, Yi-Fan Zhang et al. · 0 citations
Preprint Jul 2026

Biologically Informed Deep Neural Networks for Multi-Omic Integration, Pathway Activity Inference and Risk Stratification in Cancer

This work reports Pathway Activity Autoencoders for the multi-omics setting, which embed prior knowledge via pathway-informed architectural constraints, fostering interpretability, while preserving representational power, in the context of breast cancer.

Pedro Henrique da Costa Avelar, L. Ou-Yang, Min Wu et al. · 0 citations
Open access Aug 2026

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings

This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention and identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets.

Erik Vonkaenel, Lisa M. Bramer, Javier E. Flores et al. · 0 citations
Open access Aug 2026

A Conceptual Framework for AI-Integrated Metabolomics in Predictive Health Systems for Resource-Constrained Environments

The increasing prevalence of non-communicable diseases (NCDs) continues to place significant pressure on healthcare systems, particularly in low- and middle-income regions where access to early diagnostic infrastructure remains limited. Conventional healthcare approaches are often reactive, detecting diseases after substantial progression and reducing opportunities for timely intervention. This challenge highlights the need for predictive, affordable, and data-driven healthcare solutions that can support early diagnosis and prevention. This study proposes a conceptual framework that integrates metabolomics with artificial intelligence (AI) to support predictive health systems in resource-constrained environments. Metabolomics enables comprehensive characterization of small-molecule metabolites, providing valuable insights into physiological and pathological changes. When combined with machine learning approaches, metabolomic datasets can be analyzed to identify potential biomarkers, classify disease risks, and generate personalized healthcare insights. The proposed framework presents a multi-layered architecture consisting of metabolomic data acquisition, preprocessing, feature engineering, AI-based predictive modeling, and clinical decision-support outputs. The model emphasizes scalability through the integration of portable diagnostic technologies, cloud-based analytics, edge computing, and decentralized healthcare delivery approaches. It also considers critical implementation challenges, including data harmonization, infrastructure limitations, algorithmic bias, and ethical governance. Furthermore, the framework highlights the need for empirical validation through pilot studies, technology assessment, and multi-site evaluation to determine its feasibility, reliability, and applicability across diverse healthcare settings. By integrating biological data analysis, computational intelligence, and responsible innovation principles, this study provides a pathway toward accessible predictive and precision public health systems for underserved populations. Overall, this research contributes to the advancement of AI-enabled healthcare by proposing a scalable and context-sensitive model that bridges metabolomics, artificial intelligence, and healthcare delivery requirements in resource-constrained environments.

David Sunday Araoti · 0 citations