Skip to content
Open access

Evaluating Dataset Representativeness for Machine Learning-Based Anomaly Detection in Encrypted Network Environments

Sep 2026 · European Journal of Electrical Engineering and Computer Science · 0 citations · 13 references

Abstract

The increasing prevalence of encrypted network traffic has reduced the effectiveness of traditional intrusion detection methods that rely on payload inspection and signature-based analysis. Machine learning–based anomaly detection offers a promising alternative, but its effectiveness depends heavily on the representativeness of the datasets used for training and evaluation. Many legacy datasets were developed in unencrypted, low-volume environments and do not reflect modern encrypted traffic conditions. This study evaluated the impact of dataset representativeness on anomaly detection performance in encrypted network environments using a quantitative comparative design. Two public datasets, ISCX-VPN-NonVPN-2016 and CSE-CIC-IDS-2018, were analyzed. Logistic Regression and Random Forest classifiers were assessed using repeated stratified train-test splits with standardized preprocessing and feature engineering. Performance metrics included accuracy, precision, recall, F1-score, ROC-AUC, false positive rate, and operational latency. Results showed that feature-rich modern datasets improved detection performance and reliability, particularly by reducing false positives and enhancing classification consistency in encrypted traffic. Random Forest outperformed Logistic Regression across most metrics while maintaining acceptable latency. These findings highlight the importance of dataset representativeness in developing effective machine learning–based cybersecurity solutions for modern encrypted environments.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.