A hybrid framework integrating Recursive Feature Elimination with Cross- Validation, GridSearchCV, and Firefly Optimization for feature selection and hyperpa- rameter optimization along with SMOGN for imbalance handling is proposed.
Abstract
Software defect density prediction is essential for improving software quality and reliability; however, ac- curate prediction is challenging due to feature redundancy and imbalanced regression data. To address these issues, this study proposes a hybrid framework integrating Recursive Feature Elimination with Cross- Validation (RFECV), GridSearchCV, and Firefly Optimization (FFO) for feature selection and hyperpa- rameter optimization, along with SMOGN for imbalance handling. The framework is evaluated on six datasets using Gradient Boosting, Random Forest, Bagging, AdaBoost, Voting, and XGBoost regression models. Performance is assessed using MSE, MAE, RMSE, and SERA metrics. Experimental results show significant improvement, with RMSE reduced from 7.96 to 0.10 on the healthcare dataset and from 2.06 to 0.056 on the ANT 1.5 dataset. Comparisons with PSO, GWO, and GA further demonstrate superior accuracy, stability, and convergence behavior of FFO. Statistical validation using the Wilcoxon signed- rank test confirms the effectiveness and robustness of the proposed framework for software defect density prediction.
In smart city software systems, where interconnected services demand high reliability, Software Defect Prediction (SDP) plays a vital role and reducing maintenance costs by identifying defect-prone modules early in the Software Development Life Cycle (SDLC). Cross-Project Defect Prediction (CPDP) enables defect data fr...
Emediong Bassey Obot, Victor Anaga, Sadiq Thomas et al.· E3S Web of Conferences· 0 citations
Software defect prediction (SDP) is essential for improving software quality since it finds error-prone modules early in the development lifecycle. Current methods produce inflated and erroneous performance metrics because of data leaks, inadequate class imbalance management, and reliance on antiquated classifiers. By...
B. V. Chowdary, Sendhil Kumar B. B, D. L. Sri et al.· International Conference on...· 0 citations
Accurate software development effort estimation is essential but often hindered by high-dimensional data and the inefficiencies of handling feature selection and parameter tuning as separate, sequential processes. This study proposes an integrated Whale Optimization Algorithm–Support Vector Regression (WOA-SVR) framewo...
R. Putri, G. E. Yuliastuti, Citra Nurina Prabiantissa· JOURNAL OF APPLIED INFORMATI...· 0 citations
This study adapts TabKANet to the all-numerical, highly imbalanced SDP setting and empirically evaluates it against established baselines, using a structured ablation in order to isolate the contribution of oversampling and feature selection rather than to propose a new architecture.
Setyo Wahyu Saputro, M. Faza, Azhiman Saputra Setyo et al.· 0 citations
TabKANet is a competitive architecture for all-numerical, highly imbalanced SDP, matching strong neural baselines and surpassing TabNet, where effective class weighting alone suffices and SMOTE is counter-productive.
Muhammad Faza Azhiman Saputra, Setyo Wahyu Saputro, M. Faisal et al.· Indonesian Journal of Electr...· 0 citations
This study compared no correction, random oversampling, random undersampling, SMOTE, ADASYN, and class-weighted learning across logistic regression, decision tree, random forest, support vector machine, and neural network classifiers to find accuracy alone is unsuitable for selecting defect predictors.
L. Akpan· International Journal of App...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.