Aug 2026· 2026 12th International Conference on Big Data and Information Analytics (BigDIA)· pp. 589-596· 0 citations· 44 references
Abstract
Non-stationary data streams suffer from simultaneous data and concept drifts that degrade model generalization. Conventional Automated Machine Learning for data streams, AML4S, relies on single-pipeline architecture with univariate ADWIN detection and exhaustive post-drift reconstruction, causing knowledge waste and insufficient ensemble robustness under high-dimensional composite drift. This work proposes MDIE-AutoML, integrating PCA-aided multi-detector joint drift identification, warm-start incremental retraining, DWE-based Top-K weighted ensemble, and a drift-category-aware memory bank. Evaluation on LoanDataset and four benchmarks spanning natural, cyclic, gradual, and high-dimensional multi-class drift demonstrates that the MDIE-Top5 variant achieves 0.9311 ± 0.0097 overall accuracy on the 20,000-sample main experiment, with roughly half the standard deviation of the AML4S baseline, and ranks first in seven of nine cross-scenario tests. These results confirm that MDIE achieves superior prediction stability with manageable computational overhead across diverse non-stationary scenarios.
This work proposes DMAE, a dual-memory active ensemble learning method for multiclass imbalanced concept-drifting data streams, and proposes a composite sample-weighting formulation, PCN-Weight, which jointly models boundary difficulty, class-imbalance status, sample–prototype relations, and temporal decay to guide inc...
Meng Han, Ya-Jie Xue, Yi-Kai Li et al.· Journal of King Saud Univers...· 0 citations
Real-world datasets often exhibit evolving distributions, known as concept drift. Ignoring drift degrades predictive performance, while reliance on fixed hyperparameters further limits model adaptability under changing conditions. Adaptive learning addresses this challenge by continuously updating models online, allowi...
Mohammad Abu-Shaira, Wei-Shi Shi· Information Sciences· 0 citations
The MLIDSC introduces an automated labeling mechanism that eliminates the need for prior parameter assumptions and uses a novel weighted scheme that combines the imbalance ratio and the importance of individual instances at a given time, ensuring a focus on critical data points.
Bohnishikhan Halder, K. M. Azharul Hasan, Md. Manjur Ahmed· Applied intelligence (Boston...· 0 citations
Deployed predictive machine learning models inevitably degrade when real-world production data deviates from training distributions. Static deployment strategies cannot handle this environmental shift, causing drops in accuracy and forcing engineering teams to rely on slow, manual retraining audits. This paper presents...
Sri Charan Chowdary Konidina· International journal of dat...· 0 citations
The results demonstrate that CADEE can improve label utilization, drift adaptation, and minority class recognition in non-stationary multi-class imbalanced data streams.
Meng Han, Ya-Jie Xue, Yi-Kai Li et al.· Knowledge and Information Sy...· 0 citations