Unsupervised Domain Adaptation for Malicious Network Traffic Detection in Heterogeneous Internet Environments
Abstract
The increasing heterogeneity of network traffic and the rapid evolution of cyberattacks pose significant challenges for malicious traffic detection. Traditional intrusion detection approaches, which rely on handcrafted features and static assumptions about traffic distributions, often exhibit limited robustness when applied across different network environments. Although deep learning methods can automatically learn expressive traffic representations, their dependence on labeled data restricts their applicability in large-scale and evolving network settings. This article presents UDAMNTD, an unsupervised domain adaptation framework for malicious network traffic detection. The contribution of UDAMNTD lies in the integration of a spatial–temporal feature extraction backbone (CNN–BiLSTM with hierarchical attention) and complementary domain-alignment objectives, including adversarial alignment, maximum mean discrepancy (MMD)-based statistical matching, and reconstruction consistency. This design enables the learning of discriminative and transferable traffic representations without requiring labeled data from target domains. By jointly optimizing feature learning and domain adaptation objectives, UDAMNTD reduces distribution discrepancies between source and target domains while preserving informative traffic characteristics. The framework captures spatial–temporal traffic patterns and improves detection performance for both majority and minority attack categories under cross-domain settings. Extensive evaluations on four benchmark datasets—CICIDS2018, CIC-DDoS2019, BoT-IoT, and MAWIFlow—demonstrate that UDAMNTD consistently achieves strong empirical performance against representative unsupervised domain adaptation baselines across diverse cross-domain malicious traffic detection scenarios. These results indicate that the proposed framework provides an effective approach for learning transferable traffic representations in heterogeneous network environments.