Skip to content
Open access

A Robust Multi-Method Feature Selection Framework for Early Software Defect Prediction: Cross-Dataset Evidence from PROMISE Repository

Aug 2026 · Journal of Intelligent Decision Making and Information Science · Vol 3, pp. 1488-1506 · 0 citations · 16 references

TL;DR

The proposed multi-method feature selection framework shows good feature stability and has a high F1-score in multiple datasets with a value up to 0.42, which means it is effective at cross-dataset defect prediction, and highlights the importance of embedding imbalance-aware learning techniques.

Abstract

The software defect prediction problem is difficult because of high-dimensional feature spaces, redundancy in software metrics and severe class imbalance. There has been a large body of work proposed on machine learning methods, however, few works have focused on the stability and generalizability of selected features across heterogeneous datasets. A comprehensive multi-method feature selection framework that combines filter (Mutual Information, Chi-Square), wrapper (Recursive Feature Elimination) and embedded (Random Forest) methods and finally a consensus-based ranking mechanism for finding stable and dataset-independent software metrics is proposed. The framework is tested on several PROMISE datasets (KC1, KC2, PC1, JM1), representing various software systems and distributions of metrics. Experimental results show that the size and complexity related metrics such as Lines of Code (LOC), Halstead measures and cyclomatic complexity invariably sit on top of all data sets and selection approaches, resulting in high feature stability. The results of the classification based on Logis-tic Regression, however, have found a significant difference between the measure of precision (0.58-0.65) and the measure of re-call (0.23-0.31), which means that the classification is not so effective in detecting the minority of defective in-stances. The results show that, in imbalanced conditions, using robust feature selection is not enough for early defect prediction. The study highlights the importance of embedding imbalance-aware learning techniques (cost-sensitive learning, synthetic sampling, and ensemble models). The proposed framework shows good feature stability and has a high F1-score in multiple datasets with a value up to 0.42, which means it is effective at cross-dataset defect prediction.

Read PDF

Similar papers

Conference Aug 2026

A Hybrid Feature Selection and Ensemble Learning Framework for Software Defect Prediction with Class Imbalance Handling

Software defect prediction (SDP) is essential for improving software quality since it finds error-prone modules early in the development lifecycle. Current methods produce inflated and erroneous performance metrics because of data leaks, inadequate class imbalance management, and reliance on antiquated classifiers. By...

B. V. Chowdary, Sendhil Kumar B. B, D. L. Sri et al. · 0 citations
Open access Aug 2026

An Explainable Feature Selection and Stacking Ensemble Framework for Software Fault Prediction

Overall, the findings indicate that integrating principled feature selection with a boosting-based stacking ensemble can improve software fault prediction performance while providing greater transparency for software quality management.

Harsimran Kaur, Hardeep Singh, Amitpal Singh Sohal et al. · 0 citations
Conference Open access 2026

Cross Project Software Defect Prediction Using Machine Learning with Optimized Feature Selection

In smart city software systems, where interconnected services demand high reliability, Software Defect Prediction (SDP) plays a vital role and reducing maintenance costs by identifying defect-prone modules early in the Software Development Life Cycle (SDLC). Cross-Project Defect Prediction (CPDP) enables defect data fr...

Emediong Bassey Obot, Victor Anaga, Sadiq Thomas et al. · 0 citations
Open access Aug 2026

Research on Software Defect Prediction Based on Static and Dynamic Feature Fusion and Machine Learning

Aiming at the problems that the static-dynamic feature fusion mechanism lacks systematic multi-scenario verification, the feature-model adaptation law is unclear, and the engineering practicability of existing research conclusions is insufficient, this paper proposes a static-dynamic feature fusion defect prediction me...

Tian-Yu Yin · 0 citations
Aug 2026

Precision Software Defect Prediction Using Novel Machine Learning Approaches

This study aims to improve software defect prediction five publicly available NASA datasets by using Random Forest and Classification Network to achieve higher defect prediction accuracy compared to methods without feature selection (WOFS) and to get matrix problem the authors use Classification Network.

S. G., Santosh Santosh · 0 citations
Open access Aug 2026

Comparative Evaluation of TabKANet with Oversampling and Feature Selection Ablation for Software Defect Prediction

TabKANet is a competitive architecture for all-numerical, highly imbalanced SDP, matching strong neural baselines and surpassing TabNet, where effective class weighting alone suffices and SMOTE is counter-productive.

Muhammad Faza Azhiman Saputra, Setyo Wahyu Saputro, M. Faisal et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.