An adaptive feature selection approach to construct structured and discriminative feature subsets that remain effective across diverse datasets and learning models is proposed and shows that the performance gains are attributable to efficient, interpretable feature subsets, highlighting the importance of robustness-oriented feature selection in SDP.
Abstract
Software defect prediction (SDP) often faces challenges related to heterogeneous software metrics, classifier dependency, and severe class imbalance, which may limit the robustness and generalization of feature selection strategies. This study proposes an adaptive feature selection approach to construct structured and discriminative feature subsets that remain effective across diverse datasets and learning models. The proposed method first measures the relationship between each software metric and the defect label using absolute determination power and then applies an adaptive retention rule to iteratively retain features with stronger discriminative contribution. The evaluation was conducted on multiple public defect datasets using several classical machine learning classifiers. Unlike approaches optimized for specific classifiers, the proposed strategy emphasizes cross-classifier robustness and imbalance-aware evaluation through defect recall and Matthews correlation coefficient. Experimental results show that the proposed method achieves a competitive average MCC of 0.257 and a defect recall of 0.446 compared with baseline approaches, although the statistical tests do not indicate significant superiority. Therefore, the proposed method should be interpreted as a comparable and stable alternative for feature selection under imbalanced SDP conditions. Stability and statistical analyses further indicate that the proposed method maintains comparable performance across dataset-classifier combinations. In addition, feature compactness analysis shows that the performance gains are attributable to efficient, interpretable feature subsets, highlighting the importance of robustness-oriented feature selection in SDP.
Software defect prediction (SDP) is essential for improving software quality since it finds error-prone modules early in the development lifecycle. Current methods produce inflated and erroneous performance metrics because of data leaks, inadequate class imbalance management, and reliance on antiquated classifiers. By...
B. V. Chowdary, Sendhil Kumar B. B, D. L. Sri et al.· International Conference on...· 0 citations
The proposed multi-method feature selection framework shows good feature stability and has a high F1-score in multiple datasets with a value up to 0.42, which means it is effective at cross-dataset defect prediction, and highlights the importance of embedding imbalance-aware learning techniques.
Papiya Mukherjee, Mamta Dahiya· Journal of Intelligent Decis...· 0 citations
Overall, the findings indicate that integrating principled feature selection with a boosting-based stacking ensemble can improve software fault prediction performance while providing greater transparency for software quality management.
Harsimran Kaur, Hardeep Singh, Amitpal Singh Sohal et al.· International journal of com...· 0 citations
In smart city software systems, where interconnected services demand high reliability, Software Defect Prediction (SDP) plays a vital role and reducing maintenance costs by identifying defect-prone modules early in the Software Development Life Cycle (SDLC). Cross-Project Defect Prediction (CPDP) enables defect data fr...
Emediong Bassey Obot, Victor Anaga, Sadiq Thomas et al.· E3S Web of Conferences· 0 citations
TabKANet is a competitive architecture for all-numerical, highly imbalanced SDP, matching strong neural baselines and surpassing TabNet, where effective class weighting alone suffices and SMOTE is counter-productive.
Muhammad Faza Azhiman Saputra, Setyo Wahyu Saputro, M. Faisal et al.· Indonesian Journal of Electr...· 0 citations
This study adapts TabKANet to the all-numerical, highly imbalanced SDP setting and empirically evaluates it against established baselines, using a structured ablation in order to isolate the contribution of oversampling and feature selection rather than to propose a new architecture.
Setyo Wahyu Saputro, M. Faza, Azhiman Saputra Setyo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.