Skip to content
#protein folding Open access

Omicau Multi-Omics Benchmark Suite

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Overview A prospectively frozen benchmark suite for leakage-safe multi-omic integration. It tests predictive performance, modality utility, model-capacity effects, missingness handling, null behavior, exploratory external transport, failure handling, and compute cost without making claims of clinical utility or causal biological inference. Included datasets Synthetic controls: paired null and planted-signal families with operative structured missingness for binary classification and continuous regression. DepMap/CCLE: transcriptomics, copy number, LC-MS metabolomics, and PRMT5 dependency across 644 cell lines. TCGA BRCA: transcriptomics, copy number, and RPPA protein abundance for ductal-versus-lobular classification across 783 tumors. TCGA LGG, KIRC, and UCEC: transcriptomics and copy number for IDH status, pathological stage, and histology endpoints across 507, 507, and 500 tumors, respectively. CPTAC UCEC exploratory holdout: transcriptomics and copy number for endometrioid-versus-serous transport assessment across 95 patient-disjoint tumors. Methods and controls Nine fixed methods compare Omicau with an unmasked architecture-matched ablation, matched early and single-modality neural controls, early and single-modality linear controls, weighted late fusion, a missingness-only diagnostic, and calibrated latent partial least squares. A TCGA-UCEC complete-training-feature sensitivity tests outcome-associated technical missingness. Internal cohorts use shared group-aware partitions, training-only preprocessing, five outer folds repeated three times, 5,000 paired group bootstraps, paired DeLong tests for AUROC, 4,999 paired squared-error sign flips for R-squared, and Holm adjustment across five primary matched-capacity contrasts. Effect sizes and intervals are the primary evidence. Because repeated out-of-fold predictions share training sets, internal p-values are conditional on the frozen prediction vectors and are not unconditional population-generalization tests. Ten target permutations per cohort are coarse catastrophic-leakage diagnostics, not formal empirical tail-probability tests. TCGA and CPTAC expression scales are harmonized by a source-declared, target-blind transformation with pooled and matched-feature numerical-domain gates. The CPTAC endpoint is exploratory because it was exercised during predeposit development smoke. Literature-anchored controls remain independent of method ranking. Failed, unfavorable, discordant, non-estimable, and indeterminate outcomes remain reportable. Reproducibility The archive contains the frozen protocol, immutable source registry, download and validation code, internal and external partitions, fixed comparator implementations, statistical aggregation, schemas, environment pins, and fault-injection tests. Raw molecular matrices, participant-level data, local paths, and benchmark results are excluded. Deviations Deviation 1 - Aggregation target normalization. Final aggregation converts read-only NumPy memory-mapped target vectors to base NumPy arrays before metric and bootstrap validation while preserving scientific values and frozen randomization streams. Deviation 2 - Comparator and ablation expansion. Matched neural, unmasked, missingness-only, complete-feature, weighted late-fusion, and partial least-squares controls separate fusion value from model capacity, technical missingness, and integration strategy. All settings are fixed before definitive execution. Deviation 3 - Independent external evaluation. A CPTAC UCEC holdout adds 95 patient-disjoint assessment cases. TCGA-UCEC supplies all training and model-selection rows; shared transcriptomic and copy-number features are aligned by unique Entrez identifiers and expression scales are harmonized without using CPTAC outcomes. Deviation 4 - Primary contrast realignment. The primary contrast is Omicau minus the matched early neural control. Prior linear comparisons remain fully reported as contextual estimates and are not substituted for the matched-capacity test. Deviation 5 - External development exposure. The CPTAC endpoint was exercised during predeposit development smoke. External estimates are designated exploratory and are not treated as untouched confirmatory validation. Deviation 6 - External expression-scale correction. Predeposit development smoke exposed incompatible TCGA and CPTAC expression domains. Linear TCGA RSEM values now receive log2(x+1) after negative values are marked missing; CPTAC retains its source-declared log2 scale. Target-blind pooled and matched-feature gates validate compatibility. The correction precedes definitive execution, and earlier smoke outputs are not reused. Deviation 7 - Synthetic missingness application correction. Predeposit audit showed that registered synthetic missingness masks were not reaching model matrices. The masks now alter every synthetic method input exactly as registered. The correction precedes definitive execution, and earlier smoke outputs are not reused.

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

In the context of cloud computing, risks associated with underlying technologies, risks involving service models and outsourcing, and enterprise readiness have been recognized as potential barriers for the adoption. To accelerate cloud adoption, the concrete barriers negatively influencing the adoption decision need to be identified. Our study aims at understanding the impact of technical and security-related barriers on the organizational decision to adopt the cloud. We analyzed data collected through a web survey of 352 individuals working for enterprises consisting of decision makers as well as employees from other levels within an organization. The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability. The result from our logistic regression analysis confirms the criticality of the security concern, which results in an up to 26-fold increase in the non-adoption likelihood. Our study underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

To compete in this age of disruption, large companies cannot rely on cost efficiency, lead time reduction and quality improvement. They are now looking for ways to innovate like startups. Meanwhile, the awareness and use of the Lean startup approach have grown rapidly amongst the software startup community in recent years. This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors. A multiple case study approach is followed in the investigation. Two software product innovation projects from two large companies are examined, using a conceptual framework that is based on the method-in-action framework and extended with the previously developed Lean-Internal Corporate Venture model. Seven face-to-face in-depth interviews of the employees with different roles are conducted. Within-case analysis and cross-case comparison are applied to draw the findings from the cases. A generic process flow summarises the common key processes of Lean internal startups. The findings suggest that an internal startup that is initiated management or employees faces different challenges. A list of enablers of applying Lean startup in large companies are identified, including top management support and cross-functional team. Both cases face different inhibitors due to the different process of inception, objective of the team and type of the product. Our contributions are threefold. First, this study is one of the first attempt to investigate the use of Lean startup approach in large companies empirically. Second, the study shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context. The third is a general process of Lean internal startup and the evidence of the enablers and inhibitors of implementing it, which are both theory-informed and empirically grounded.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

Affects--emotions and moods--have an impact on cognitive processing activities and the working performance of individuals. It has been established that software development tasks are undertaken through cognitive processing activities. Therefore, we have proposed to employ psychology theory and measurements in software engineering (SE) research. We have called it "psychoempirical software engineering". However, we found out that existing SE research has often fallen into misconceptions about the affect of developers, lacking in background theory and how to successfully employ psychological measurements in studies. The contribution of this paper is threefold. (1) It highlights the challenges to conduct proper affect-related studies with psychology; (2) it provides a comprehensive literature review in affect theory; and (3) it proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

It is essential for startups to quickly experiment business ideas by building tangible prototypes and collecting user feedback on them. As prototyping is an inevitable part of learning for early stage software startups, how fast startups can learn depends on how fast they can prototype. Despite of the importance, there is a lack of research about prototyping in software startups. In this study, we aimed at understanding what are factors influencing different types of prototyping activities. We conducted a multiple case study on twenty European software startups. The results are two folds; firstly we propose a prototype-centric learning model in early stage software startups. Secondly, we identify factors occur as barriers but also facilitators for prototyping in early stage software startups. The factors are grouped into (1) artifacts, (2) team competence, (3) collaboration, (4) customer and (5) process dimensions. To speed up a startup’s progress at the early stage, it is important to incorporate the learning objective into a well-defined collaborative approach of prototyping.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#human-computer interacti... Open access May 2025

Linker-free PROTACs efficiently induce the degradation of oncoproteins

Proteolysis-targeting chimeras (PROTACs) present a potentially effective strategy against various diseases via selective proteolysis. How to increase the efficacy of PROTACs remains challenging. Here, we explore the necessity of the linker, which has been deemed as an integral part of heterobifunctional PROTACs. Adopting single amino acid-based degradation signals, we find that the linker is not a required feature of the PROTACs. Notably, the linker-free PROTAC, Pro-BA, exhibits superior efficacy over its linker-bearing counterparts in degrading EML4-ALK and inhibiting lung cancer cell growth, as Pro-BA induces a stronger interaction between the target and the E3 ubiquitin ligase. Pro-BA is a water-soluble, orally administered degrader that significantly inhibits the tumor growth in a xenograft mouse model. The broad applicability of this linker-free PROTAC strategy is further validated through the development of BCR-ABL degrader. Our study introduces a design paradigm for PROTACs, potentially facilitating the advancement of more efficient therapeutic degraders. Linkers are traditionally seen as important for PROTAC activity. Here, the authors demonstrate that linker-free PROTACs can outperform traditional designs, marking a paradigm shift in PROTAC development for targeted protein degradation.

Jianchao Zhang, Congli Chen, Xiao Chen et al. · 41 citations
#machine learning Open access Nov 2025

mRNABERT: advancing mRNA sequence design with a universal language model and comprehensive dataset

Designing effective mRNA sequences for therapeutics remains a formidable challenge. Inspired by successes in protein design, language models (LMs) are now being applied to RNA, but progress is often impeded by the lack of comprehensive training data. Existing models are frequently limited to UTR or CDS regions, restricting their application for complete mRNA sequences. We introduce mRNABERT, a robust, all-in-one mRNA designer pre-trained on the largest available mRNA dataset. To enhance performance, we propose a dual tokenization scheme with a cross-modality contrastive learning framework to integrate semantic information from protein sequences. On a comprehensive benchmark, mRNABERT demonstrates state-of-the-art performance, outperforming previous models in the majority of tasks for 5’ UTR and CDS design, RNA-binding protein (RBP) site prediction, and full-length mRNA property prediction. It also surpasses large protein models in several related tasks. In conclusion, mRNABERT’s superior performance across these diverse tasks signifies a substantial leap forward in mRNA research and therapeutic development. Designing complete mRNA sequences for new vaccines and therapies is a complex challenge. Here, the authors develop mRNABERT, a foundational AI model that designs entire mRNA sequences and demonstrates superior performance across comprehensive benchmarks.

Ying Xiong, Aowen Wang, Yu Kang et al. · 22 citations · ⚡1

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.