The findings indicate that feature importance should be validated separately for each target dataset and that successful cross-dataset deployment is likely to require explicit domain adaptation, even when datasets share the same feature-extraction schema.
Abstract
Static malware detection with machine learning relies critically on which features are extracted from Portable Executable (PE) binaries. Most prior work evaluates feature selection on a single benchmark, leaving open whether features generalize across heterogeneous datasets. We investigate cross-dataset feature discriminability by quantifying the consistency of important features across heterogeneous malware datasets. We introduce Cross-Validated Feature Ranking (CVFR), a lightweight wrapper that aggregates Random Forest (RF) feature importance across repeated cross-validation runs to attach per-feature confidence intervals and a stability score (SNR) to an otherwise standard ranking. CVFR is not a new feature-selection algorithm and is not claimed to improve accuracy; it is positioned as the uncertainty-quantification instrument that enables the cross-dataset discriminability study. RF, Multilayer Perceptron (MLP), LightGBM, and One-Class SVM are evaluated under a two-level stratified cross-validation protocol (25 repeated evaluations) with within-fold SMOTE on four PE-header benchmark datasets: EMBER, BIG15, Malicia, and BODMAS. RF achieves the highest F1 and AUC-ROC on every dataset (mean 98.72% F1, 0.998 AUC), consistently outperforming MLP across all 25 fold-seed pairs (Wilcoxon signed-rank test, $p \lt 0.0001$ ), and remaining best or tied against a competitive LightGBM gradient-boosting baseline. An ablation study confirms that CVFR achieves predictive performance comparable to single-run impurity ranking and permutation importance, while additionally providing uncertainty estimates for feature importance at lower computational cost than permutation importance. Jaccard analysis reveals limited feature overlap (7.9–27.6%) among independently collected datasets, while datasets sharing the same extraction schema (EMBER and BODMAS) reach 83.8% overlap. Transfer experiments show that the high-overlap pair (EMBER $\leftrightarrow $ BODMAS, 83.8% Jaccard) experiences larger F1 degradation (20–36 pp) than the low-overlap pair (BIG $15\leftrightarrow $ Malicia, 8–13 pp), suggesting that feature-space overlap alone is insufficient to predict transferability and that temporal distributional shift may exert a stronger influence on cross-dataset generalization. Because the transfer analysis covers only two dataset pairs (two seeds each) that co-vary simultaneously in feature overlap, temporal collection gap, and class balance, this transfer observation is hypothesis-generating rather than statistically conclusive. Subject to that caveat, the findings indicate that feature importance should be validated separately for each target dataset and that successful cross-dataset deployment is likely to require explicit domain adaptation, even when datasets share the same feature-extraction schema.
The increasing sophistication of modern malware, particularly in the form of polymorphic and packed variants, poses significant challenges to traditional detection systems. While machine learning-based approaches have shown promise, many existing methods rely on single feature representations and are evaluated using me...
Akshaya Mogili, G. Lishita, R. N· 2026 International Conferenc...· 0 citations
Detecting and classifying Android malware families remains challenging due to high feature dimensionality, class imbalance, and the high cost of expert-labeled data. Semi-supervised learning (SSL) offers a way to leverage unlabeled samples, but prior works rarely test whether SSL benefits generalize across classifier t...
Static malware family classification can use evidence from raw byte content, disassembly, and Portable Executable metadata. Each representation captures different characteristics of a malware sample. We propose a multi-view stacking framework for Microsoft BIG 2015 that represents each sample through seven static featu...
Kim-Huu Tran, Nguyen-Huy Le-Huu, Quoc-Huy Nguyen et al.· International Conference on...· 0 citations
Android malware is growing rapidly, making detection more challenging and necessitating systems that are both accurate and efficient. This study aims to design a lightweight and reliable framework for Android malware detection and classification that reduces computational costs while maintaining strong performance. The...
V. Dwivedi, A. Waoo· Review of Computer Engineeri...· 0 citations
The results confirm that hybrid resampling combined with optimized gradient boosting improves classification reliability, especially in addressing severe class imbalance and enhancing recognition capability across diverse Android malware families.
Ali Nur Ikhsan, Adam Prayogo Kuncoro, Debby Ummul Hidayah et al.· Journal of Information Syste...· 0 citations
The diversity of today’s malware and the conflicting criteria for predictive performance, computational efficiency, dependability, and interpretability have made the choice of a suitable malware detection model more complicated. Existing research focuses predominantly on predictive performance, while the multidimension...
Husam Jasim Mohammed, Riyadh Rahef Nuiaa Alogaili, Mohanad S. Jabbar et al.· Mathematical and Computation...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.