Skip to content
Open access

Integrating Opcode N-Grams and Word Embeddings for Enhanced Malware Classification: A Comparative Study with Transformer-Based Representations

Sep 2026 · Electronics · 0 citations · 15 references

Abstract

This work proposes a comparative framework for malware classification that evaluates the synergy between traditional feature engineering and modern deep learning architectures. Our methodology follows two primary paths: first, we integrate opcode n-grams with word-embedding techniques (Word2Vec, Doc2Vec, and FastText) to capture local execution patterns in dense vector spaces. Second, we evaluate end-to-end representations using transformer-based models (BERT and ViT) and a raw opcode-based 1D Convolutional Neural Network (1D-CNN) to determine if effective features can be learned without explicit n-gram engineering. Both pathways are rigorously tested across a suite of classifiers, including Support Vector Machine (SVM), Random Forest (RF), k-Nearest Neighbor (k-NN), and CNNs. Experimental results for multi-class classification demonstrate that while transformer-based models offer high automated feature extraction capabilities, the combination of opcode n-grams with word embeddings remains a highly effective and interpretable approach for detecting real-world malware.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.