Skip to content

Diagonal Multi-omics Integration of Heterogenous Datasets

Aug 2026 · 0 citations · 22 references
Mathematics Computer Science

TL;DR

A novel characteristic of dataset heterogeneity is introduced by employing the norm of the difference between the maximum and minimum points in the classical terms of functional analysis.

Abstract

In this paper, we consider methods for the diagonal multi-omics integration of heterogeneous datasets. Several approaches to the nature of biological heterogeneity are analyzed and developed to comprehend more clearly the generated differences. Specifically, the extremal trace problems for the coupled Laplacian on sets homeomorphic to the Stiefel manifold embedded in the complex Euclidean space are investigated. The gradient ascent method for the maximization problem is elaborated in the classical terms of functional analysis, which is of significant interest in itself. On this basis, we introduce a novel characteristic of dataset heterogeneity by employing the norm of the difference between the maximum and minimum points.

View source

Similar papers

Open access Aug 2026

Multi-Omic Spectral Clustering with the Flag Mean

One common goal in multi-omics studies is to identify subgroups within the study’s cohort. Many methods create subgroups through unsupervised clustering, and, to our knowledge, all of these methods attempt to infer a common subspace or latent cluster across multiple views. We argue that this is a strict assumption that may not truly exist in many datasets. We employ the classic spectral clustering algorithm in conjunction with the flag manifold. Together, this allows for differing cluster structures across the omics profiles, leading to a novel approach for more flexible subtyping in multi-omics studies. We study a data set on ventilator associated pneumonia in children. These data contain airway microbiome and transcriptome. Through simulation studies, we demonstrate that the flag mean of separate clustering subspaces can accurately capture the span of a joint clustering space. It is also robust to varying noise structures and number of features across omics profiles. Our proposed method also outperforms popular multi-omics clustering methods in the presence of differing group sizes. This proposed method is general enough to apply to other multi-omics studies as well as any multi-view study that uses spectral clustering. Code for these methods are written in R and are freely available through GitHub at https://github.com/Ghoshlab/MMOC or through CRAN at https://cran.r-project.org/web/packages/MMOC/index.html

Charlie M. Carpenter, Ziwei Tian, J. Harris et al. · 0 citations
Open access Feb 2025

Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data

A fundamental design pattern in biomolecular studies is to assay the same set of samples (organisms, tissue biopsies, or individual cells) by multiple different ‘omics assays. Group Factor Analysis (GFA) and its adaptation to high-dimensional settings, Multi-Omics Factor Analysis (MOFA), are widely used as a first-line approach to analyse such data and are effective in detecting patterns of correlation, organize them into so-called latent factors, and identify common and assay-specific factors. However, in many applications a subset of the found factors just rediscovers already known covariates (e.g., disease subtypes, environmental covariates) while others may represent genuine novelty. Here, we present Semi-supervised Omics Factor Analysis (SOFA), a method that incorporates known covariates into the model upfront and focuses the factor discovery on novel sources of variation. We show SOFA’s effectiveness for discovering novel patterns by applying it to cancer, brain development and heart failure multi-omic data sets.

Tümay Capraz, Harald Vöhringer, Klaus Sebastian Augusto Kruger Serrano et al. · 2 citations
Preprint Aug 2026

Knowledge-guided Pattern Discovery via Coupled Tensor Factorizations

In order to understand complex systems such as the human metabolome or human brain, different sensing technologies are used, generating complex data. These datasets are often multiway, i.e., with more than two axes of variation such as a subjects by metabolites by time array. While tensor factorizations have successfully revealed interpretable patterns from such complex data, they have so far been mainly data-driven. On the other hand, there is more to data -- there are computational models (of these systems), which are rich sources of prior information. In this paper, we introduce a knowledge-guided approach that brings together data and computational models by jointly analyzing real data and simulated data (generated using a computational model) using coupled tensor factorizations with linear coupling. Our experiments on real metabolomics measurements demonstrate that guiding the analysis of such noisy data with simulated data improves the pattern discovery performance while also revealing potential discrepancies between data and computational models.

Gaute Johannessen, G. Ploeg, E. Acar · 0 citations

Related blog posts