Skip to content
Conference

A Hybrid Feature Selection and Ensemble Learning Framework for Software Defect Prediction with Class Imbalance Handling

Aug 2026 · International Conference on Information Security and Cryptology · pp. 614-620 · 0 citations · 17 references

Abstract

Software defect prediction (SDP) is essential for improving software quality since it finds error-prone modules early in the development lifecycle. Current methods produce inflated and erroneous performance metrics because of data leaks, inadequate class imbalance management, and reliance on antiquated classifiers. By integrating a soft voting ensemble of six contemporary gradient boosting classifiers XGBoost, LightGBM, CatBoost, Random Forest, Extra Trees, and Gradient Boosting with Mutual Information filtering, this study provides a MI-based feature selection technique that gets beyond these restrictions. To avoid data leakage, SMOTE is only used on training data to address class imbalance. Five publicly accessible PROMISE repository datasets CM1, PC1, KC1, KC2, and JM1 are used in the experiments. For all five datasets, the average accuracy is 82.86%, while the average F1-score and ROC-AUC are 46.92% and 0.778, respectively. Previous studies on the same datasets only give accuracy numbers of up to 94% for imbalanced defect datasets, when the number of non-faulty modules significantly exceeds that of defective ones. This is a deceptive figure that hides insufficient ability to detect defects. F1-score, Precision, Recall, and ROC-AUC are used to measure performance; these metrics offer a more accurate evaluation than accuracy alone. The results show that MI-based feature selection improves prediction performance while lowering computational overhead; on all five datasets, contemporary ensemble techniques regularly beat traditional classifiers.

View source

Similar papers

Open access 2023

COMPARATIVE EVALUATION OF CLASS-IMBALANCE CORRECTION TECHNIQUES FOR SOFTWARE DEFECT PREDICTION

This study compared no correction, random oversampling, random undersampling, SMOTE, ADASYN, and class-weighted learning across logistic regression, decision tree, random forest, support vector machine, and neural network classifiers to find accuracy alone is unsuitable for selecting defect predictors.

L. Akpan · 0 citations
Conference Open access 2026

Cross Project Software Defect Prediction Using Machine Learning with Optimized Feature Selection

In smart city software systems, where interconnected services demand high reliability, Software Defect Prediction (SDP) plays a vital role and reducing maintenance costs by identifying defect-prone modules early in the Software Development Life Cycle (SDLC). Cross-Project Defect Prediction (CPDP) enables defect data fr...

Emediong Bassey Obot, Victor Anaga, Sadiq Thomas et al. · 0 citations
Open access Aug 2026

An Explainable Feature Selection and Stacking Ensemble Framework for Software Fault Prediction

Overall, the findings indicate that integrating principled feature selection with a boosting-based stacking ensemble can improve software fault prediction performance while providing greater transparency for software quality management.

Harsimran Kaur, Hardeep Singh, Amitpal Singh Sohal et al. · 0 citations
Open access Aug 2026

Firefly Optimization-Based Feature Selection for Software Defect Density Prediction

A hybrid framework integrating Recursive Feature Elimination with Cross- Validation, GridSearchCV, and Firefly Optimization for feature selection and hyperpa- rameter optimization along with SMOGN for imbalance handling is proposed.

Jasmeet Kaur, Arvinder Kaur, Kamaldeep Kaur · 0 citations
Open access Aug 2026

Comparative Evaluation of TabKANet with Oversampling and Feature Selection Ablation for Software Defect Prediction

TabKANet is a competitive architecture for all-numerical, highly imbalanced SDP, matching strong neural baselines and surpassing TabNet, where effective class weighting alone suffices and SMOTE is counter-productive.

Muhammad Faza Azhiman Saputra, Setyo Wahyu Saputro, M. Faisal et al. · 0 citations
Open access Sep 2026

Multiclass Defect Classification from Legacy Foundry Data: A Decision Support System for Reducing Manual Inspection Time

A machine learning-based decision support system for multiclass defect detection, utilizing exclusively heterogeneous legacy process data to minimize manual inspection time in foundries, and a Naive Bayes Stacking meta-classifier effectively neutralizes single-algorithm inductive biases.

Joachim Denker, Loui Al-Shrouf, M. Jelali · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.