Comparative analysis of Feature Selection Algorithms and Performances on Medical Classification Problem
TL;DR
The study recommends adopting CAE as the default choice in routine clinical applications due to its efficiency, and using RFE in research studies requiring maximum accuracy, to guiding researchers and clinicians in selecting the optimal combination of feature selection and classification algorithms according to their specific medical context.
Abstract
Machine learning-based medical classification has faced significant challenges related to the high dimensionality of medical data, leading to increased computational complexity, overfitting, and poor clinical interpretability of models. This study aims to evaluate the performance of five feature selection algorithms available within the WEKA platform CAE, GRE, IGE, ORE, and RFE algorithms in improving the accuracy of four classification models: Random Forest, Naïve Bayes, KNN, and SVM, across six medical datasets: heart attacks, diabetes, breast cancer, liver disorders, and hepatitis. Experiments perform using 10-fold cross-validation, and Accuracy, Sensitivity, Specificity, F1-score, AUC-ROC, and Time complexity (Tc) is calculated. The results show that applying Feature Selection algorithms led to an average accuracy improvement of 0.360%, from 80.20% to 80.56%. The effect is most pronounced with the SVM classifier, which improved by 0.62 percentage points. The RFE and CAE algorithms combined with SVM achieved the highest overall accuracy 92.4%, with RFE offering the best balance between 80.56% accuracy and a computational time of 0.28 seconds. The results also showed that the optimal selection algorithm varied depending on the nature of the dataset; RFE is optimal for breast cancer and heart attacks, ORE for hepatitis, and CAE for diabetes. The study recommends adopting CAE as the default choice in routine clinical applications due to its efficiency, and using RFE in research studies requiring maximum accuracy. These findings contribute to guiding researchers and clinicians in selecting the optimal combination of feature selection and classification algorithms according to their specific medical context.