The findings indicate that alignment parameter selection substantially influences CPDP performance and that a fixed global default is inadequate.
Abstract
In cross-project defect prediction (CPDP), divergence in feature distribution between the source and target often fails when predicting defects across unrelated projects. Such a mismatch is typical rather than exceptional in real deployment scenarios. To address this, the Transfer Component Analysis (TCA) technique projects the source and target into a shared subspace, where Maximum Mean Discrepancy is minimized. However, combining fixed ensemble classifiers with TCA on NASA datasets is yet to be explored. Previous studies either searched ensemble compositions adaptively or conflated alignment with source selection. In this study, we trained a two-layer hybrid ensemble of Bagging and AdaBoost classifiers, alongside a Logistic Regression meta-learner, on TCA-aligned features across 20 directed source–target pairs. We used five PROMISE datasets, each having 21 McCabe and Halstead features. Experiments were conducted, and the findings show that, against an unaligned baseline using the same ensemble, TCA alignment increased mean AUC by 0.131 (0.625 to 0.755), mean F1 by 0.155, and mean MCC by 0.125. Wilcoxon signed-rank tests confirmed significance across all three metrics (p < 0.003, rank-biserial r = 0.714, Cliff’s Delta d ≥ 0.545). A full ablation showed that alignment was the dominant contributor to performance, with TCA improving AUC by +0.131 over the unaligned baseline. In contrast, SMOTE traded a small AUC reduction (−0.014) for substantial F1 gains (+0.134), while the stacking layer provided modest improvements in F1 and MCC. TCA reduced MMD across all 20 source–target pairs by a mean of 83.3%. Sensitivity analysis further showed that the TCA subspace dimensionality parameter, k, exhibited non-monotone, pair-specific AUC sensitivity, with optimal values ranging from 5 to 30. These findings indicate that alignment parameter selection substantially influences CPDP performance and that a fixed global default is inadequate.
TLSS_ICFS mitigates source-target distribution divergence, overcomes the limitations of traditional CPDP methods, and provides a stable, high-performance solution for cross-project defect prediction in data-scarce scenarios.
H. Fu, Ran Mo, Xin-Ya Mu et al.· International Conference on...· 0 citations
Overall, the findings indicate that integrating principled feature selection with a boosting-based stacking ensemble can improve software fault prediction performance while providing greater transparency for software quality management.
Harsimran Kaur, Hardeep Singh, Amitpal Singh Sohal et al.· International journal of com...· 0 citations
This study adapts TabKANet to the all-numerical, highly imbalanced SDP setting and empirically evaluates it against established baselines, using a structured ablation in order to isolate the contribution of oversampling and feature selection rather than to propose a new architecture.
Setyo Wahyu Saputro, M. Faza, Azhiman Saputra Setyo et al.· 0 citations
Software defect prediction (SDP) is essential for improving software quality since it finds error-prone modules early in the development lifecycle. Current methods produce inflated and erroneous performance metrics because of data leaks, inadequate class imbalance management, and reliance on antiquated classifiers. By...
B. V. Chowdary, Sendhil Kumar B. B, D. L. Sri et al.· International Conference on...· 0 citations
TabKANet is a competitive architecture for all-numerical, highly imbalanced SDP, matching strong neural baselines and surpassing TabNet, where effective class weighting alone suffices and SMOTE is counter-productive.
Muhammad Faza Azhiman Saputra, Setyo Wahyu Saputro, M. Faisal et al.· Indonesian Journal of Electr...· 0 citations
This study compared no correction, random oversampling, random undersampling, SMOTE, ADASYN, and class-weighted learning across logistic regression, decision tree, random forest, support vector machine, and neural network classifiers to find accuracy alone is unsuitable for selecting defect predictors.
L. Akpan· International Journal of App...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.