Reliable Machine Learning for Omics Data
This PhD thesis investigates the reliable application of machine learning (ML) and deep learning (DL) methods to omics data, with particular emphasis on high-dimensional, low-sample-size settings commonly encountered in foodomics and agronomy. The work focuses on improving methodological rigor, interpretability, and reproducibility in applied omics research. The dissertation is organized into two main parts. The first part addresses methodological aspects of ML for omics data, including evaluation strategies and hybrid modeling approaches. In particular, it examines the interaction between cross-validation and early stopping in neural network training, identifying common pitfalls such as information leakage and biased performance estimation. Furthermore, the thesis explores hybrid neural network architectures that integrate mechanistic domain knowledge into data-driven learning, framing the problem as a multi-objective optimization task that balances predictive accuracy with mechanistic consistency. The second part focuses on applied case studies in foodomics and agronomy. It presents robust and explainable deep learning models for SNP-based phenotype prediction, demonstrating statistically significant improvements in predictive performance through adaptive optimization, regularization, and data augmentation strategies. The work also employs SHAP-based explainability methods to identify biologically relevant features and ensure transparent interpretation of model predictions. In addition, the thesis introduces a research-stage MLOps framework for organizing complex omics machine learning workflows, improving experiment traceability, reproducibility, and comparability. Overall, the dissertation contributes to the development of more reliable, interpretable, and reproducible ML methodologies for omics data analysis by combining methodological innovation with practical applications in foodomics and agronomy.