Comparative evaluation of resampling techniques for improving the performance of classification algorithms on a breast cancer dataset
Abstract
Class imbalance is a common problem in machine learning classification tasks and may negatively affect the reliability and predictive performance of classification models. The research object is the comprehensive evaluation of different resampling techniques applied to a publicly available breast cancer dataset consisting of 569 records categorized as benign and malignant. To address class imbalance, four resampling techniques available in the WEKA (Waikato Environment for Knowledge Analysis) environment, namely Resample, SMOTE (Synthetic Minority Over-sampling Technique), SpreadSubSample, and StratifiedRemoveFolds, were applied to the dataset. The resulting datasets were evaluated using NB (Naive Bayes), KNN (K-Nearest Neighbors), and DT (Decision Tree) classification algorithms. Algorithm performance was assessed using accuracy, precision, recall, and F1-score metrics. The experimental results indicate that resampling techniques have a significant impact on classification performance. Among the evaluated methods, StratifiedRemoveFolds produced the highest accuracy values, achieving 0.982 with Naive Bayes and 0.964 with Decision Tree. In contrast, SpreadSubSample generally reduced the performance of all evaluated algorithms. Furthermore, the findings revealed that different resampling techniques influence algorithms in different ways. These results demonstrate that the effectiveness of a resampling method depends on the characteristics of both the dataset and the classification algorithm used. The research highlights the practical application of appropriate resampling strategies in medical decision-support systems and other machine learning applications involving imbalanced datasets, where reliable classification performance is required