Hybrid Representation Learning for Robust Multiclass Malware Classification Under Class Imbalance
Abstract
The increasing sophistication of modern malware, particularly in the form of polymorphic and packed variants, poses significant challenges to traditional detection systems. While machine learning-based approaches have shown promise, many existing methods rely on single feature representations and are evaluated using metrics that do not fully capture performance under class imbalance. This paper presents a hybrid representation learning framework for multiclass Windows malware classification that integrates byte-level structural features with opcode-level semantic features. Specifically, byte histogram representations are combined with TF-IDF weighted opcode sequences to capture complementary aspects of executable behavior. To improve robustness, a soft-voting ensemble of Logistic Regression, LightGBM, and Multilayer Perceptron models is employed. The framework is evaluated on a subset of 2,500 samples (9 malware families) from the Microsoft Malware Classification Challenge dataset under naturally imbalanced class distributions. LightGBM was the strongest individual model, reaching 98.8% accuracy and a macro-F1 of 0.98 on a held-out 500 -sample test split; the soft-voting ensemble of Logistic Regression, LightGBM, and MLP achieved 97.6% accuracy (macro-F1 0.96), offering more balanced precision/recall across models at a small cost in peak accuracy. The study further highlights the importance of macro-level evaluation metrics and feature complementarity in achieving reliable malware classification. The findings suggest that hybrid feature representations, when combined with ensemble learning and imbalance-aware training strategies, offer a practical and scalable direction for real-world malware detection systems.