ANOMALY DETECTION ALGORITHMS IN STREAMING DATA: COMPARATIVE ANALYSIS AND APPLICATION
Abstract
This paper presents a systematic review and comparative analysis of state-of-the-art anomaly detection algorithms for streaming data - one of the fundamental problems in data mining, cybersecurity, and industrial monitoring. Unlike existing surveys, which mostly address the batch setting, all methods here are examined from a streaming perspective: per-observation processing cost, resident memory footprint, and the ability to update the model incrementally. An explicit set of comparison criteria is introduced - computational complexity, memory requirements, adaptivity, robustness to concept drift, and interpretability - together with the scales and rules used to assign qualitative ratings. Four major algorithm classes are examined: statistical methods (Z-score, EWMA, CUSUM, Shewhart control charts), distance- and density-based methods (k-NN, LOF, ILOF), ensemble methods - Isolation Forest and its streaming adaptations Half-Space Trees, iForestASD, and Robust Random Cut Forest - and neural network architectures including autoencoders (Autoencoder, VAE, LSTM-Autoencoder) and transformer-based models (Anomaly Transformer, TranAD). For each class, both strengths and intrinsic limitations are systematised: the dependence of statistical methods on distributional assumptions, the degradation of distance-based methods in high-dimensional spaces, the sensitivity of Isolation Forest to complex data structures and masking effects, and the high resource demands and limited interpretability of neural models. Concept drift adaptation strategies are discussed: sliding window, weighted learning, and the ADWIN algorithm. Evaluation results on real-world datasets are described - KDD Cup 1999 (cybersecurity) and SMAP NASA (industrial IoT) - along with the well-documented methodological limitations of these benchmarks and evaluation protocols, which imply that published metric values should be treated as upper bounds. Key practical application domains are considered: network intrusion detection, predictive equipment maintenance, financial transaction monitoring, and medical diagnostics. The results demonstrate that algorithm selection is governed by a trade-off between computational efficiency, accuracy, and adaptability; practical selection guidelines are formulated for five typical deployment scenarios. Promising directions for future research are identified: federated learning, explainable models (SHAP, LIME), transformer-based architectures, and hardware acceleration on FPGAs.