Skip to content
Open access

An Explainable Multi-Criteria Decision-Making Framework for Evaluating Malware Detection Models Across Heterogeneous Datasets

Sep 2026 · Mathematical and Computational Applications · 0 citations · 30 references

Abstract

The diversity of today’s malware and the conflicting criteria for predictive performance, computational efficiency, dependability, and interpretability have made the choice of a suitable malware detection model more complicated. Existing research focuses predominantly on predictive performance, while the multidimensional decision process required for practical model selection remains insufficiently addressed. To address this gap, this study provides an explainability-aware hybrid multi-criteria decision-making (MCDM) framework that systematically evaluates and ranks malware detection models across heterogeneous malware datasets. The methodology incorporates predictive performance, computational efficiency, false positive rate, and a composite Explainability Index into a single decision procedure. The Explainability Index integrates explanation stability, sparsity, and expert relevance, enabling interpretability to be explicitly considered in the model-selection process. The hybrid criteria weights are obtained by combining the Analytic Hierarchy Process (AHP) with the entropy weighting method, thereby integrating expert-driven criterion importance with data-driven variability. The final ranking is obtained using the Technique for Order Preference by Similarity to the Ideal Solution (TOPSIS). Four candidate detection models, including Random Forest, XGBoost, Convolutional Neural Network, and Long Short-Term Memory, are independently used to validate the framework on the CIC-MalMem-2022 and CICMalDroid2020 datasets. Under a harmonized evaluation protocol without assuming direct cross-dataset predictive transfer, the experimental results rank XGBoost first with a TOPSIS closeness score of 0.670, followed by Random Forest with 0.624. Sensitivity and ablation analyses further show that the model ranking remains stable while changes in criterion weighting and framework components produce measurable variations in the multi-criteria preference structure. Overall, the framework provides a transparent and multidimensional alternative to conventional performance-centered malware model evaluation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.