Skip to content
Open access

Enhancing Cross-Project Defect Prediction via Transfer Component Analysis and Hybrid Ensembles

Aug 2026 · Software · 0 citations · 27 references

TL;DR

The findings indicate that alignment parameter selection substantially influences CPDP performance and that a fixed global default is inadequate.

Abstract

In cross-project defect prediction (CPDP), divergence in feature distribution between the source and target often fails when predicting defects across unrelated projects. Such a mismatch is typical rather than exceptional in real deployment scenarios. To address this, the Transfer Component Analysis (TCA) technique projects the source and target into a shared subspace, where Maximum Mean Discrepancy is minimized. However, combining fixed ensemble classifiers with TCA on NASA datasets is yet to be explored. Previous studies either searched ensemble compositions adaptively or conflated alignment with source selection. In this study, we trained a two-layer hybrid ensemble of Bagging and AdaBoost classifiers, alongside a Logistic Regression meta-learner, on TCA-aligned features across 20 directed source–target pairs. We used five PROMISE datasets, each having 21 McCabe and Halstead features. Experiments were conducted, and the findings show that, against an unaligned baseline using the same ensemble, TCA alignment increased mean AUC by 0.131 (0.625 to 0.755), mean F1 by 0.155, and mean MCC by 0.125. Wilcoxon signed-rank tests confirmed significance across all three metrics (p < 0.003, rank-biserial r = 0.714, Cliff’s Delta d ≥ 0.545). A full ablation showed that alignment was the dominant contributor to performance, with TCA improving AUC by +0.131 over the unaligned baseline. In contrast, SMOTE traded a small AUC reduction (−0.014) for substantial F1 gains (+0.134), while the stacking layer provided modest improvements in F1 and MCC. TCA reduced MMD across all 20 source–target pairs by a mean of 83.3%. Sensitivity analysis further showed that the TCA subspace dimensionality parameter, k, exhibited non-monotone, pair-specific AUC sensitivity, with optimal values ranging from 5 to 30. These findings indicate that alignment parameter selection substantially influences CPDP performance and that a fixed global default is inadequate.

Read PDF

Similar papers

Sep 2026

TLSS_ICFS: A novel two-stage cross-project defect prediction approach

TLSS_ICFS mitigates source-target distribution divergence, overcomes the limitations of traditional CPDP methods, and provides a stable, high-performance solution for cross-project defect prediction in data-scarce scenarios.

H. Fu, Ran Mo, Xin-Ya Mu et al. · 0 citations
Open access Aug 2026

An Explainable Feature Selection and Stacking Ensemble Framework for Software Fault Prediction

Overall, the findings indicate that integrating principled feature selection with a boosting-based stacking ensemble can improve software fault prediction performance while providing greater transparency for software quality management.

Harsimran Kaur, Hardeep Singh, Amitpal Singh Sohal et al. · 0 citations

TabKANet for Software Defect Prediction: An Oversampling and Feature-Selection Ablation Study

This study adapts TabKANet to the all-numerical, highly imbalanced SDP setting and empirically evaluates it against established baselines, using a structured ablation in order to isolate the contribution of oversampling and feature selection rather than to propose a new architecture.

Setyo Wahyu Saputro, M. Faza, Azhiman Saputra Setyo et al. · 0 citations
Conference Aug 2026

A Hybrid Feature Selection and Ensemble Learning Framework for Software Defect Prediction with Class Imbalance Handling

Software defect prediction (SDP) is essential for improving software quality since it finds error-prone modules early in the development lifecycle. Current methods produce inflated and erroneous performance metrics because of data leaks, inadequate class imbalance management, and reliance on antiquated classifiers. By...

B. V. Chowdary, Sendhil Kumar B. B, D. L. Sri et al. · 0 citations
Open access Aug 2026

Comparative Evaluation of TabKANet with Oversampling and Feature Selection Ablation for Software Defect Prediction

TabKANet is a competitive architecture for all-numerical, highly imbalanced SDP, matching strong neural baselines and surpassing TabNet, where effective class weighting alone suffices and SMOTE is counter-productive.

Muhammad Faza Azhiman Saputra, Setyo Wahyu Saputro, M. Faisal et al. · 0 citations
Open access 2023

COMPARATIVE EVALUATION OF CLASS-IMBALANCE CORRECTION TECHNIQUES FOR SOFTWARE DEFECT PREDICTION

This study compared no correction, random oversampling, random undersampling, SMOTE, ADASYN, and class-weighted learning across logistic regression, decision tree, random forest, support vector machine, and neural network classifiers to find accuracy alone is unsuitable for selecting defect predictors.

L. Akpan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.