Skip to content
Open access

MalBERT-Temporal: Transformer-Based Zero-Day Malware Detection in Windows Executable Binaries Under Strict Temporal Isolation

2026 · International Journal of Advanced Computer Science and Applications · 0 citations · 29 references

TL;DR

A Transformer-based malware detection approach named MalBERT-Temporal that reshapes the 2,381-dimensional BODMAS PE feature vector into 16 contiguous feature-group tokens and then processes these tokens with Transformer encoder layers employing multi-head self-attention to model interactions among all tokens is introduced.

Abstract

In current malware detection benchmarks, random train/test splits are commonly used, allowing temporal leakage to occur and obscuring performance degradation caused by concept drift. Moreover, traditional classifiers that operate on a one-dimensional PE feature vector do not explicitly model long-range interactions between structurally distant feature groups. This study introduces a Transformer-based malware detection approach named MalBERT-Temporal that reshapes the 2,381-dimensional BODMAS PE feature vector into 16 contiguous feature-group tokens and then processes these tokens with Transformer encoder layers employing multi-head self-attention to model interactions among all tokens. The proposed approach is evaluated using a strict temporal protocol in which the training data comprise only pre-2020 samples, and the test data span the complete 2020 evaluation timeline. Each model is independently calibrated using a fixed validation set with a false-positive-rate budget of 0.1% and is evaluated monthly and on malware families unseen during training. At the calibrated security threshold, MalBERT-Temporal achieves an F1-score of 97.84% with a false-positive rate of 0.138%, outperforming a 1-D CNN baseline, which achieves an F1-score of 92.01% and a false-positive rate of 3.388%, and a Random Forest baseline, which achieves an F1-score of 67.29%. Welch’s t-tests confirm statistically significant differences in monthly F1-scores between the proposed approach and both baselines. Moreover, MalBERT-Temporal maintains monthly F1-scores within a 2.2-point band throughout the nine-month evaluation period, indicating improved robustness to temporal distribution shift. A component-wise ablation traces the improvement to the tokenized self-attention mechanism itself: substituting a token-wise feed-forward block for self-attention while keeping every other component costs 2.73 F1 points, and removing the tokenization costs 1.85 points, while positional embeddings, the CLS token, multi-scale pooling, encoder depth, and nonlinearity are each worth 0.47 points or less. A matched random-split control that leaves the model, the preprocessing, the training budget, and the calibration procedure unchanged gives an F1-score of 98.86%, which confirms that chronological evaluation is the stricter protocol.

Read PDF

Similar papers

Review Open access Aug 2026

A Comparative Evaluation of Malware Families and Machine-Learning Detection Techniques, and an Optimized Stacked-Ensemble Model for Predicting Software Maliciousness

An efficient stacked-ensemble model that estimates how likely a given executable is to be malicious and is compared against recent malware research/types are compared and identified for future research work/area.

Deepak Singh Rana, Sushil Chandra Dimri · 0 citations
Open access Aug 2026

Deep Learning for Malware Detection: TransformerBased Analysis of Windows Executables

A Transformerbased deep learning approach for detecting malicious Windows Portable Executable files is presented, significantly outperforming baseline methods including LightGBM, MalConv, and LSTM.

E. Baghirov · 0 citations
Conference Aug 2026

DeepGuard: Explainable Behavioral Malware Detection Using Hybrid CNN-BiLSTM-Attention Deep Learning on Dynamic API Call Sequences

Recent rapid increases in sophisticated malware have severely challenged traditional signature-based security systems, particularly in their ability to recognize zero-day and polymorphic threats. This paper introduces “DeepGuard,” an AI/Machine-Learning-powered behavioral malware detection framework built around the an...

Vidya Gavekar, Amar Anant Shinde, S. Lodha et al. · 0 citations
Open access Aug 2026

Real-Time Detection and Mitigation of Prompt Injection Attacks in LLM-Integrated Enterprise Systems

Large language models (LLMs) embedded in enterprise workflows cannot structurally distinguish legitimate instructions from adversarial ones in the same token stream, making prompt injection OWASP's top LLM risk for two consecutive editions a persistent threat across direct and indirect vectors. This paper presents Prom...

Fatimah Alhamzawi · 0 citations
Open access Aug 2026

Empirical Evaluation and New Insights of Concept Drift in ML-Based Android Malware Detection

Despite outstanding results, machine learning-based Android malware detection models struggle with concept drift, where rapidly evolving malware characteristics degrade model effectiveness. This study examines the impact of concept drift on Android malware detection, evaluating two datasets and nine machine learning an...

Ahmed Sabbah, Mohammed F. Kharma, Radi Jarrar et al. · 0 citations
#machine learning Preprint Aug 2026

REPLICANT: Learning Policies for Evading and Hardening Malware Detectors

This work presents Replicant, a deep reinforcement learning framework that learns the realistic task of evasion under a strict label-only black-box threat model and demonstrates that learning the task of evasion not only results in stronger attack performance but provides a better signal for hardening malware detectors...

Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.