An Explainable Multi-Criteria Decision-Making Framework for Evaluating Malware Detection Models Across Heterogeneous Datasets
Abstract
The diversity of today’s malware and the conflicting criteria for predictive performance, computational efficiency, dependability, and interpretability have made the choice of a suitable malware detection model more complicated. Existing research focuses predominantly on predictive performance, while the multidimensional decision process required for practical model selection remains insufficiently addressed. To address this gap, this study provides an explainability-aware hybrid multi-criteria decision-making (MCDM) framework that systematically evaluates and ranks malware detection models across heterogeneous malware datasets. The methodology incorporates predictive performance, computational efficiency, false positive rate, and a composite Explainability Index into a single decision procedure. The Explainability Index integrates explanation stability, sparsity, and expert relevance, enabling interpretability to be explicitly considered in the model-selection process. The hybrid criteria weights are obtained by combining the Analytic Hierarchy Process (AHP) with the entropy weighting method, thereby integrating expert-driven criterion importance with data-driven variability. The final ranking is obtained using the Technique for Order Preference by Similarity to the Ideal Solution (TOPSIS). Four candidate detection models, including Random Forest, XGBoost, Convolutional Neural Network, and Long Short-Term Memory, are independently used to validate the framework on the CIC-MalMem-2022 and CICMalDroid2020 datasets. Under a harmonized evaluation protocol without assuming direct cross-dataset predictive transfer, the experimental results rank XGBoost first with a TOPSIS closeness score of 0.670, followed by Random Forest with 0.624. Sensitivity and ablation analyses further show that the model ranking remains stable while changes in criterion weighting and framework components produce measurable variations in the multi-criteria preference structure. Overall, the framework provides a transparent and multidimensional alternative to conventional performance-centered malware model evaluation.