Anomaly Detection of Water Physicochemical Indicators Based on an Isolation Forest–Random Forest Fusion Algorithm
Abstract
To overcome the challenges of limited abnormal samples, diverse anomaly patterns, and the insufficient capability of conventional threshold-based methods in detecting multivariate water quality abnormalities, this study develops an anomaly detection approach based on an Isolation Forest–Random Forest fusion algorithm. Hourly automatic monitoring data from an urban river section during 2021–2023 were used to construct a multivariate dataset covering thermal, acid–base, oxygen-related, optical, ionic, nutrient, and organic pollution characteristics. In the proposed framework, standardized water quality samples were first processed by the Isolation Forest model for unsupervised anomaly screening, through which anomaly scores were obtained. These scores were then introduced as enhanced features and combined with the original physicochemical variables as the input of the Random Forest model for anomaly reclassification and key factor identification. Experimental results show that the proposed IF-RF fusion model performed better than conventional threshold-based detection and several representative machine learning methods, demonstrating stronger overall performance across multiple evaluation metrics. Specifically, the F1-score and AUC of the fusion model reached 0.902 and 0.962, respectively. The feature importance results indicate that the anomaly score generated by Isolation Forest contributed most to the final classification, while dissolved oxygen, turbidity, ammonia nitrogen, and COD were identified as the dominant physicochemical factors affecting anomaly detection. Overall, this study provides a practical data-driven method for automatic water quality monitoring, anomaly warning, and refined management of urban rivers.