Encrypted Network Traffic Classification Using Generative and Contrastive Self-Supervised Learning
Abstract
Network traffic classification is crucial for various applications, encompassing network provisioning, malware detection, and resource management. In contemporary networks, the prevalence of encrypted protocols presents a challenge to existing classification techniques. Deep learning has exhibited promising results in encrypted traffic classification; however, its reliance on extensive labeled data impedes its efficacy in real-world scenarios where labeled data is scarce. This paper explores self-supervised learning, specifically generative and contrastive learning, as a solution to this challenge. While self-supervised learning is well-established in computer vision and natural language processing, its application to sequential data like network traffic remains unexplored. The paper utilizes two self-supervised approaches, generative and contrastive, tailored to achieve high accuracy in classifying encrypted network traffic with minimal labeled data. Evaluation across four publicly available datasets shows that the proposed self-supervised approaches outperform corresponding supervised baselines, an existing self-supervised method, and two state-of-the-art few-shot learning approaches by approximately 3%, 11%, and 17%, respectively in terms of accuracy. In addition, the proposed approaches exhibit transferability, showcasing the ability to apply learned knowledge to different datasets not seen during training. This paper advances the understanding and application of self-supervised learning in the domain of network traffic classification, offering viable solutions for scenarios with limited labeled data.