Skip to content
Open access

Comparative analysis of Feature Selection Algorithms and Performances on Medical Classification Problem

2026 · Libyan Journal of Medical and Applied Sciences · 0 citations

TL;DR

The study recommends adopting CAE as the default choice in routine clinical applications due to its efficiency, and using RFE in research studies requiring maximum accuracy, to guiding researchers and clinicians in selecting the optimal combination of feature selection and classification algorithms according to their specific medical context.

Abstract

Machine learning-based medical classification has faced significant challenges related to the high dimensionality of medical data, leading to increased computational complexity, overfitting, and poor clinical interpretability of models. This study aims to evaluate the performance of five feature selection algorithms available within the WEKA platform CAE, GRE, IGE, ORE, and RFE algorithms in improving the accuracy of four classification models: Random Forest, Naïve Bayes, KNN, and SVM, across six medical datasets: heart attacks, diabetes, breast cancer, liver disorders, and hepatitis. Experiments perform using 10-fold cross-validation, and Accuracy, Sensitivity, Specificity, F1-score, AUC-ROC, and Time complexity (Tc) is calculated. The results show that applying Feature Selection algorithms led to an average accuracy improvement of 0.360%, from 80.20% to 80.56%. The effect is most pronounced with the SVM classifier, which improved by 0.62 percentage points. The RFE and CAE algorithms combined with SVM achieved the highest overall accuracy 92.4%, with RFE offering the best balance between 80.56% accuracy and a computational time of 0.28 seconds. The results also showed that the optimal selection algorithm varied depending on the nature of the dataset; RFE is optimal for breast cancer and heart attacks, ORE for hepatitis, and CAE for diabetes. The study recommends adopting CAE as the default choice in routine clinical applications due to its efficiency, and using RFE in research studies requiring maximum accuracy. These findings contribute to guiding researchers and clinicians in selecting the optimal combination of feature selection and classification algorithms according to their specific medical context.

Read PDF

Similar papers

Open access Sep 2026

Intelligent Predictive Analytics using Machine Learning: A Comparative Evaluation of Classification Algorithms for High-Dimensional Data

Overall, the findings identify XGBoost as the most effective algorithm for reliable predictive analytics in high-dimensional data environments, while demonstrating that model selection should balance predictive performance with computational efficiency.

Muhammad Awais, Muhammad Haad, Hameed Hussain et al. · 0 citations
Open access Sep 2026

A Comparative Study of Machine Learning Algorithms for Detecting Heart Disease

Cardiovascular diseases remain a leading cause of mortality worldwide. Identifying underlying clinical phenotypes early, such as distinct categories of chest pain, is vital for diagnostic triage and downstream medical decision-making. This study evaluates the performance of four prominent machine learning algorithms –...

Twana Abdulqader Mohammed, S. Salh · 0 citations
Preprint Aug 2026

Transforming Heart Disease Prediction with Advanced Machine Learning Techniques

The research work concludes that ML models, when properly tuned and validated, can significantly assist in the early diagnosis of heart disease, offering critical support for clinical decision-making.

Sami Ullah, Muhammad Mohsin Khan · 0 citations
Open access Sep 2026

High-Dimensional Breast Cancer Classification: Evaluating Trade-offs Between Accuracy, Robustness, and Computational Cost

Background: Breast cancer is the leading cause of cancer-related mortality in females, with 2.3 million new cases diagnosed annually. Machine learning (ML) algorithms have the potential to improve diagnostic accuracy through the analysis of high-dimensional data. However, the lack of standardized benchmarks across diff...

Ali A. Hamad, Naaman Omar · 0 citations
Open access 2026

A Comparative Analysis of Machine Learning Algorithms for the Early Prediction of Diabetes with an Evaluation of Class-Imbalance Handling

The study comes to the conclusion that headline accuracy is an unreliable guide in imbalanced medical prediction, that imbalance handling can change a model's practical usefulness, and that this benefit is strongly algorithm-dependent, meaning that the decision to resample should be based on the algorithm and the scree...

A. Oduroye, Temilade Opanuga, Esther Tosin Akanbi et al. · 0 citations
Open access Sep 2026

Interpretable Heart Disease Prediction: Optimizing Machine Learning Models via Metaheuristic Ivy Algorithm

Cardiovascular diseases remain a major global health burden, making the development of accurate and interpretable prediction models important for clinical decision support. In this study, the metaheuristic Ivy Algorithm was employed to perform two-stage hyperparameter optimization for five machine learning classifiers,...

Yang Jiang, Zi-Hao Zuo, Rui Liang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.