A variational autoencoder attention fusion model for zero-day exploit behavior detection VAE-AttNet
Abstract
Zero-day exploit (ZDE) attacks are among the most severe threats to critical infrastructures such as cloud computing, industrial control, smart grids, and connected vehicles, owing to their unknown, stealthy, and highly destructive nature. Traditional signature- and rule-based intrusion detection is largely ineffective against them, while existing deep learning methods struggle with global distribution modeling, local contextual dependencies, and low-sample generalization on high-dimensional, sparse, temporally correlated traffic and system-call data. This paper proposes VAE-AttNet, which fuses a Variational Autoencoder (VAE) with Multi-Head Self-Attention. Unlike prior hybrid designs that apply attention to raw features, VAE-AttNet confines self-attention to the VAE’s low-dimensional latent space and couples it with a joint anomaly score fusing reconstruction error and attention entropy. A VAE encoder maps behavioral features into a latent space via evidence-lower-bound maximization; self-attention then models context over the latent sequence, and a joint reconstruction–attention-entropy score with a dynamic 99th-percentile threshold identifies zero-day attacks. On CICIDS2017, UNSW-NB15, and ADFA-LD, VAE-AttNet reaches 97.83% accuracy, 96.42% F1, and 0.9871 AUC on CICIDS2017 — 2.69%–17.57% above eight baselines (Autoencoder, Isolation Forest, One-Class SVM, LSTM-AE, etc.) in relative terms, including a 2.59–3.30 percentage-point absolute gain over the strongest baseline LSTM-AE — with an average 8.56% F1 gain over LSTM-AE across six cross-dataset settings and F1 = 0.9023 using only 10% of training data. These results are scoped to the three evaluated datasets under a hold-out-attack-type protocol; validation on production-scale, adversarial traffic remains future work.