Skip to content
Open access

Cross-Dataset Feature Discriminability in Static Malware Detection: A CVFR-Based Study Across Four Heterogeneous Benchmarks

2026 · IEEE Access · Vol 14, pp. 115424-115439 · 0 citations · 27 references
Computer Science

TL;DR

The findings indicate that feature importance should be validated separately for each target dataset and that successful cross-dataset deployment is likely to require explicit domain adaptation, even when datasets share the same feature-extraction schema.

Abstract

Static malware detection with machine learning relies critically on which features are extracted from Portable Executable (PE) binaries. Most prior work evaluates feature selection on a single benchmark, leaving open whether features generalize across heterogeneous datasets. We investigate cross-dataset feature discriminability by quantifying the consistency of important features across heterogeneous malware datasets. We introduce Cross-Validated Feature Ranking (CVFR), a lightweight wrapper that aggregates Random Forest (RF) feature importance across repeated cross-validation runs to attach per-feature confidence intervals and a stability score (SNR) to an otherwise standard ranking. CVFR is not a new feature-selection algorithm and is not claimed to improve accuracy; it is positioned as the uncertainty-quantification instrument that enables the cross-dataset discriminability study. RF, Multilayer Perceptron (MLP), LightGBM, and One-Class SVM are evaluated under a two-level stratified cross-validation protocol (25 repeated evaluations) with within-fold SMOTE on four PE-header benchmark datasets: EMBER, BIG15, Malicia, and BODMAS. RF achieves the highest F1 and AUC-ROC on every dataset (mean 98.72% F1, 0.998 AUC), consistently outperforming MLP across all 25 fold-seed pairs (Wilcoxon signed-rank test, $p \lt 0.0001$ ), and remaining best or tied against a competitive LightGBM gradient-boosting baseline. An ablation study confirms that CVFR achieves predictive performance comparable to single-run impurity ranking and permutation importance, while additionally providing uncertainty estimates for feature importance at lower computational cost than permutation importance. Jaccard analysis reveals limited feature overlap (7.9–27.6%) among independently collected datasets, while datasets sharing the same extraction schema (EMBER and BODMAS) reach 83.8% overlap. Transfer experiments show that the high-overlap pair (EMBER $\leftrightarrow $ BODMAS, 83.8% Jaccard) experiences larger F1 degradation (20–36 pp) than the low-overlap pair (BIG $15\leftrightarrow $ Malicia, 8–13 pp), suggesting that feature-space overlap alone is insufficient to predict transferability and that temporal distributional shift may exert a stronger influence on cross-dataset generalization. Because the transfer analysis covers only two dataset pairs (two seeds each) that co-vary simultaneously in feature overlap, temporal collection gap, and class balance, this transfer observation is hypothesis-generating rather than statistically conclusive. Subject to that caveat, the findings indicate that feature importance should be validated separately for each target dataset and that successful cross-dataset deployment is likely to require explicit domain adaptation, even when datasets share the same feature-extraction schema.

Read PDF

Similar papers

Conference Aug 2026

Hybrid Representation Learning for Robust Multiclass Malware Classification Under Class Imbalance

The increasing sophistication of modern malware, particularly in the form of polymorphic and packed variants, poses significant challenges to traditional detection systems. While machine learning-based approaches have shown promise, many existing methods rely on single feature representations and are evaluated using me...

Akshaya Mogili, G. Lishita, R. N · 0 citations
#machine learning Preprint Aug 2026

Classifier-Dependent Benefits of Pseudo-Labeling for Semi-Supervised Android Malware Attribution

Detecting and classifying Android malware families remains challenging due to high feature dimensionality, class imbalance, and the high cost of expert-labeled data. Semi-supervised learning (SSL) offers a way to leverage unlabeled samples, but prior works rarely test whether SSL benefits generalize across classifier t...

Md Rafid Islam, Zahid Hasan, Hafizah Rahman · 0 citations
Conference Aug 2026

Disagreement-Aware Multi-View Stacking for Robust Malware Classification

Static malware family classification can use evidence from raw byte content, disassembly, and Portable Executable metadata. Each representation captures different characteristics of a malware sample. We propose a multi-view stacking framework for Microsoft BIG 2015 that represents each sample through seven static featu...

Kim-Huu Tran, Nguyen-Huy Le-Huu, Quoc-Huy Nguyen et al. · 0 citations
Open access Sep 2026

A lightweight static-analysis framework with optimized feature selection for android malware detection and classification

Android malware is growing rapidly, making detection more challenging and necessitating systems that are both accurate and efficient. This study aims to design a lightweight and reliable framework for Android malware detection and classification that reduces computational costs while maintaining strong performance. The...

V. Dwivedi, A. Waoo · 0 citations
Open access Aug 2026

Optimizing Multiclass Android Malware Family Classification Using SMOTE-Tomek Links and XGBoost

The results confirm that hybrid resampling combined with optimized gradient boosting improves classification reliability, especially in addressing severe class imbalance and enhancing recognition capability across diverse Android malware families.

Ali Nur Ikhsan, Adam Prayogo Kuncoro, Debby Ummul Hidayah et al. · 0 citations
Open access Sep 2026

An Explainable Multi-Criteria Decision-Making Framework for Evaluating Malware Detection Models Across Heterogeneous Datasets

The diversity of today’s malware and the conflicting criteria for predictive performance, computational efficiency, dependability, and interpretability have made the choice of a suitable malware detection model more complicated. Existing research focuses predominantly on predictive performance, while the multidimension...

Husam Jasim Mohammed, Riyadh Rahef Nuiaa Alogaili, Mohanad S. Jabbar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.