Research on Software Defect Prediction Based on Static and Dynamic Feature Fusion and Machine Learning
Abstract
Aiming at the problems that the static-dynamic feature fusion mechanism lacks systematic multi-scenario verification, the feature-model adaptation law is unclear, and the engineering practicability of existing research conclusions is insufficient, this paper proposes a static-dynamic feature fusion defect prediction method based on multi-model comparison. Taking Equinox and Eclipse open-source Java project datasets from Kaggle platform as research objects, this paper constructs a feature system including 18-dimensional static CK-OO metrics and 20-dimensional dynamic change features, compares the prediction performance of Logistic Regression, Random Forest, XGBoost and Stacking ensemble model under static, dynamic and fusion feature sets. 5-fold cross-validation is used to ensure experimental reliability, and F1-Score is taken as the core evaluation index to carry out controlled experiments. The results show that on the balanced Equinox dataset, Random Forest with fusion features achieves the best single-model performance with an F1-Score of 0.7619 and Stacking model achieves 11.83% performance improvement in the static feature scenario; on the imbalanced Eclipse dataset, XGBoost with static features achieves the best performance with an F1-Score of 0.6512; CVS version control entropy, weighted method complexity, and class response are cross-scenario general core prediction features. This study reveals the core influence mechanism of dataset category balance and project scale on feature-model adaptability, and provides an empirical reference for defect prediction practice of different types of software projects.