Time series data with and without anomalies from a continuous distillation plant
Abstract
Reliable detection of process anomalies remains a challenge in industrial chemical plants. The ability of machine learning (ML) to recognize patterns has triggered numerous research efforts to apply ML to anomaly detection (AD). Simulation-based benchmarks, such as the Tennessee Eastman process, are widely used to develop and train AD methods. Well-annotated real-world process data, which are crucial for meaningful research advancements, are lacking due to industry-wide proprietary limitations. To overcome this issue, we present an openly accessible dataset of time series generated from an industry-like continuous distillation mini-plant under constant operating conditions. The data generated have different complexities: water runs, a heteroazeotropic separation of n-butanol–water, and a reactive process to produce a fuel additive. Chemical systems, plant setup, and anomalies (whether encountered or manually induced) are described alongside the process data. The complete dataset, including sensor and actuator data, annotations marking anomalies, and other metadata, is available in open access via an online Zenodo repository. It serves as training and testing data for ML-based AD and other data-driven applications.