Data-Centric Machine Learning for Reliable Industrial Systems: Approaches to Data-Centric Challenges in Industrial ML
The use of Machine Learning (ML) is rapidly expanding across diverse scientific and engineering domains. ML offers a powerful advantage over traditional modeling approaches for predictive modeling and analysis of variables of interest. This makes it particularly useful for developing advanced methods in analytical chemistry and residential energy systems, including forecasting (such as predicting hot water demand or chromatographic peak behavior), data quality assessment (such as detecting sensor drift or anomalous consumption patterns), and fault detection (such as identifying heat pump malfunctions or degraded separation performance). While traditional modeling approaches struggle to fully exploit complex, high-dimensional features, existing ML studies in the target domains often rely on limited datasets and lack automated and adaptive frameworks capable of handling the scale, variability, and non-stationary data generated in real operational settings. The availability of large amounts of data in this digital era offers unprecedented opportunities for analysis. However, the successful application of ML depends critically on the quality, preparation, and robustness of ML models and their underlying data. In industrial systems, these requirements are shaped by several interacting factors, particularly feature representation, model robustness, and adaptation under changing conditions. The challenges addressed in this research include data quality management, model selection, robustness assessment, and adaptation under changing real-world conditions. The main contributions of this research are: (i) a semi-automatic data preparation workflow with domain-specific feature engineering for large-scale oligonucleotide chromatography datasets; (ii) an unsupervised quality-centric evaluation framework that automatically clusters input data by quality level without requiring labeled annotations; (iii) the FIUL-Data fault injection framework, which quantifies the resilience boundaries of ML models under controlled data degradation; and (iv) a composite adaptive framework that integrates predictive ML with anomaly detection to enable demand-driven heat pump management in residential energy systems. Together, these contributions demonstrate that reliable industrial ML is achieved not by increasing model complexity, but through systematic data-centric practices including structured data preparation, quality-aware pipelines, robustness testing, and adaptive learning, applied across two complementary industrial domains.