Skip to content
Open access

COMPARATIVE EVALUATION OF CLASS-IMBALANCE CORRECTION TECHNIQUES FOR SOFTWARE DEFECT PREDICTION

2023 · International Journal of Applied Science and Engineering Review · 0 citations · 23 references

TL;DR

This study compared no correction, random oversampling, random undersampling, SMOTE, ADASYN, and class-weighted learning across logistic regression, decision tree, random forest, support vector machine, and neural network classifiers to find accuracy alone is unsuitable for selecting defect predictors.

Abstract

Class imbalance can make software defect predictors appear successful while missing defective modules. This study compared no correction, random oversampling, random undersampling, SMOTE, ADASYN, and class-weighted learning across logistic regression, decision tree, random forest, support vector machine, and neural network classifiers. KC1 and PC1 NASA/PROMISE datasets (3,218 modules; 403 defective) were evaluated by stratified five-fold cross-validation. Imputation, scaling, and correction were fitted only within training folds. Precision, sensitivity, specificity, F1-score, balanced accuracy, ROCAUC, PR-AUC, and confusion matrices were reported. Across classifiers, baseline balanced accuracy was 0.580; corrected means ranged from 0.696 to 0.720. Correction increased sensitivity but generally reduced precision and specificity. ROS achieved the highest mean F1-score (0.388), while SMOTE achieved the highest mean PR-AUC (0.378). A Friedman comparison indicated heterogeneity among techniques, followed by Holm-adjusted paired Wilcoxon tests. No approach dominated every classifier or metric. Leakage-safe correction and multi-metric assessment are essential; accuracy alone is unsuitable for selecting defect predictors.

Read PDF

Similar papers

Conference Aug 2026

A Hybrid Feature Selection and Ensemble Learning Framework for Software Defect Prediction with Class Imbalance Handling

Software defect prediction (SDP) is essential for improving software quality since it finds error-prone modules early in the development lifecycle. Current methods produce inflated and erroneous performance metrics because of data leaks, inadequate class imbalance management, and reliance on antiquated classifiers. By...

B. V. Chowdary, Sendhil Kumar B. B, D. L. Sri et al. · 0 citations
Review Aug 2026

Automated software debugging and bug prediction through the use of machine learning and deep learning

The results indicate that traditional ML models, especially random forest and extra trees, are still very effective for metric-based defect prediction, while DL and multi-modal approaches need to be fed with richer software artifacts to reach their full potential.

Amro Mohammad Abed Alfattah Abdin, Mohanad Alayedi, Ahmad M. Jaradat · 0 citations
Open access Aug 2026

Comparative Evaluation of TabKANet with Oversampling and Feature Selection Ablation for Software Defect Prediction

TabKANet is a competitive architecture for all-numerical, highly imbalanced SDP, matching strong neural baselines and surpassing TabNet, where effective class weighting alone suffices and SMOTE is counter-productive.

Muhammad Faza Azhiman Saputra, Setyo Wahyu Saputro, M. Faisal et al. · 0 citations

TabKANet for Software Defect Prediction: An Oversampling and Feature-Selection Ablation Study

This study adapts TabKANet to the all-numerical, highly imbalanced SDP setting and empirically evaluates it against established baselines, using a structured ablation in order to isolate the contribution of oversampling and feature selection rather than to propose a new architecture.

Setyo Wahyu Saputro, M. Faza, Azhiman Saputra Setyo et al. · 0 citations
Open access Sep 2026

Multiclass Defect Classification from Legacy Foundry Data: A Decision Support System for Reducing Manual Inspection Time

A machine learning-based decision support system for multiclass defect detection, utilizing exclusively heterogeneous legacy process data to minimize manual inspection time in foundries, and a Naive Bayes Stacking meta-classifier effectively neutralizes single-algorithm inductive biases.

Joachim Denker, Loui Al-Shrouf, M. Jelali · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.