Skip to content
Open access

Data-Centric Machine Learning for Reliable Industrial Systems: Approaches to Data-Centric Challenges in Industrial ML

TL;DR

This research demonstrates that reliable industrial ML is achieved not by increasing model complexity, but through systematic data-centric practices including structured data preparation, quality-aware pipelines, robustness testing, and adaptive learning, applied across two complementary industrial domains.

Abstract

The use of Machine Learning (ML) is rapidly expanding across diverse scientific and engineering domains. ML offers a powerful advantage over traditional modeling approaches for predictive modeling and analysis of variables of interest. This makes it particularly useful for developing advanced methods in analytical chemistry and residential energy systems, including forecasting (such as predicting hot water demand or chromatographic peak behavior), data quality assessment (such as detecting sensor drift or anomalous consumption patterns), and fault detection (such as identifying heat pump malfunctions or degraded separation performance). While traditional modeling approaches struggle to fully exploit complex, high-dimensional features, existing ML studies in the target domains often rely on limited datasets and lack automated and adaptive frameworks capable of handling the scale, variability, and non-stationary data generated in real operational settings. The availability of large amounts of data in this digital era offers unprecedented opportunities for analysis. However, the successful application of ML depends critically on the quality, preparation, and robustness of ML models and their underlying data. In industrial systems, these requirements are shaped by several interacting factors, particularly feature representation, model robustness, and adaptation under changing conditions. The challenges addressed in this research include data quality management, model selection, robustness assessment, and adaptation under changing real-world conditions. The main contributions of this research are: (i) a semi-automatic data preparation workflow with domain-specific feature engineering for large-scale oligonucleotide chromatography datasets; (ii) an unsupervised quality-centric evaluation framework that automatically clusters input data by quality level without requiring labeled annotations; (iii) the FIUL-Data fault injection framework, which quantifies the resilience boundaries of ML models under controlled data degradation; and (iv) a composite adaptive framework that integrates predictive ML with anomaly detection to enable demand-driven heat pump management in residential energy systems. Together, these contributions demonstrate that reliable industrial ML is achieved not by increasing model complexity, but through systematic data-centric practices including structured data preparation, quality-aware pipelines, robustness testing, and adaptive learning, applied across two complementary industrial domains.

Read PDF

Similar papers

Open access 2020

Data-Centric Engineering: Integrating Simulation, Machine Learning, and Statistics

This article delves deep into the confluence of simulation, ML, and statistics, showcasing how they synergize to improve engineering workflows and emphasizes that DCE is not just a technological advancement but a foundational strategy for next-generation engineering solutions.

Benjamin Scott · 0 citations
Open access 2019

Predictive Maintenance in Industry 4.0 Using Machine Learning Techniques

A structured methodology for ML-based PdM frameworks, covering data-driven, physics-based, and hybrid approaches, including supervised, unsupervised, and deep learning models is proposed, offering valuable insights for developing efficient and scalable PdM solutions.

Sithik Shah · 0 citations
Sep 2026

Cluster Indices for Real‐Time Monitoring

As technology advances, the volume, variety, and velocity of data generation continue to grow, leading to the emergence of big data analytics, which aims to extract valuable insights from these extensive datasets. The increasing volume of data presents several opportunities and challenges. In the context of Statistic...

Unknown authors · 0 citations
Conference Jul 2026

AI-based Analytical System for Anomaly Detection in Industrial Processes

This study develops and evaluates an AI-based analytical system for detecting anomalies in industrial processes. The work reviews major sources of risk in industrial control systems, distinguishes point, contextual, and collective anomalies, and summarizes the principal machine-learning approaches used for industrial a...

Mehdiyeva Almaz, Ahmedov Elmar, Uzakov Gulom et al. · 0 citations
Review Sep 2026

A State-Of-The-Art Review of Industrial Time Series Data Analysis: Methods and Applications

The Industrial Internet of Things (IIoT) is characterized by the generation of vast amounts of time-series data. Modern IIoT systems enable efficient collection, storage, and querying of massive industrial time-series data, making the processing and analysis of such data a key enabler for data-driven decision-making in...

Li-Lan Liu, Yixiang Zhang, Yanning Sun et al. · 0 citations
Review Aug 2026

Graph Machine Learning: An Opportunity for Power Systems

Modern power systems face growing operational complexity driven by the integration of renewable energy sources, decentralization, and the need for real-time decision-making across a wide range of timescales. Addressing these challenges traditionally relies on model-based methods that, while accurate, can be too slow fo...

Martin Sadric, Sebastian Pütz, Christian Nauck et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.