A comparative analysis of machine learning and deep learning models for early oral cancer detection using multi-risk epidemiological data from multiple countries
Jul 2026· African Journal Of Applied Research· Vol 12, pp. 395-412· 1 citation· 23 references
TL;DR
The results showed that early prediction of oral cancer using non-diagnostic epidemiological and behavioural features alone is difficult, which means that epidemiological data should be combined with clinical and biomarker-based data to develop more accurate early diagnosis systems.
Abstract
Purpose: This study aims to evaluate the effectiveness of machine learning, deep learning, and ensemble learning methods for the early diagnosis of oral cancer using multi-risk epidemiological and behavioural data collected from several countries. The main goal is to examine whether non-diagnostic factors can support early prediction before clear clinical signs appear.
Design/Methodology/Approach: The study used a dataset containing 84,922 records, including pre-diagnosis and post-diagnosis variables. Post-diagnosis variables were used only for interpretation and analysis to avoid data leakage. Several models were tested, including Random Forest, Logistic Regression, Support Vector Machine, XGBoost, LightGBM, TabNet, MLP classifier, and a voting-based ensemble model. The models were evaluated using accuracy, recall, precision, F1-score, and ROC-AUC.
Research Limitation: The main limitation is that the study did not use clinical diagnostic features, medical imaging, genetic markers, or laboratory biomarkers, which may improve prediction performance.
Findings: The results showed that when only non-diagnostic epidemiological and behavioural features were used, all models achieved performance close to random classification. This means that early prediction of oral cancer using these features alone is difficult. The study also showed clear regional and economic differences among countries in oral cancer prevalence, feature importance, treatment costs, and productivity losses.
Practical Implication: The findings suggest that epidemiological data should be combined with clinical and biomarker-based data to develop more accurate early diagnosis systems.
Social Implication: The study highlights the need for better awareness, early screening programs, and improved access to diagnostic services, especially in developing countries.
Originality/Value: This study provides a realistic evaluation of oral cancer early prediction using non-diagnostic data and emphasises the importance of avoiding data leakage in medical AI studies.
Proper diagnosis and prognosis of colon cancer is essential to enhancing patient outcomes and informing individual treatment plans. This paper has formulated a deep learning model that is powered by the R-MobileNet design to perform effective colon cancer classification through clinical structured data. The model relied on a variety of data, featuring systematic pre-processing measures and careful validation to achieve as much predictive reliability as possible. Measures of evaluation related to ROC curve analysis, classification accuracy, predicted class distribution, and confusion matrix were evaluated, and the R-MobileNet model consistently performed well on validation datasets, with an area under the curve (AUC) of 0.75 and a high classification accuracy of 0.96. Predicted classes distribution and confusion matrix indicate successful discrimination of the relevant cell types and highlight the ability of this method to accurately detect and stratify prognosis. These findings confirm the incorporation of advanced deep learning algorithms such as R-MobileNet into the clinical decision-making process in managing colon cancer.
A. S.· International Conference on...· 0 citations
Cervical cancer remains one of the leading causes of cancer-related mortality among women worldwide, particularly in low- and middle-income countries where access to early screening and diagnosis is limited. Accurate prediction of cervical cancer threat using machine learning techniques can support early intervention, improve patient outcomes and supporting clinical decision-making. This study presents a comparative analysis of five widely used supervised machine learning algorithms such as decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), Logistic Regression (LR), and XGBoost or cervical cancer risk prediction. A publicly available cervical cancer dataset was preprocessed to address missing values, class imbalance, and feature scaling where necessary. The models were trained and evaluated using standard performance metrics namely; Accuracy, Precision, Recall, F1-score, and Area Under the Receiver Operating Characteristic Curve (AUC-ROC). Experimental results demonstrate that XGBoost and Random Forest achieved superior predictive performance due to their robustness against overfitting and ability to model complex feature interactions, but XGBoost is better than Random Forest with a little difference. Logistic Regression provided a strong baseline with high interpretability, while SVM exhibited competitive classification performance depending on parameter optimization. Decision Tree offered transparent decision rules but was more susceptible to overfitting compared to the other models. The findings highlight the effectiveness of each algorithm and emphasize the potential of machine learning models in facilitating early identification of individuals at high risk of cervical cancer. This comparative study provides valuable insights for researchers and healthcare practitioners seeking to develop accurate and interpretable predictive models for cervical cancer screening, facilitating early detection and intervention, reducing the burden of cervical cancer.
Agwu Joy Nneka, Ituma Chinagolum· International Journal of Sci...· 0 citations
Based on the medical statistical analyses of relevant information, in recent years, breast cancer is one of the leading causes of cancer-related deaths in women worldwide. Hence, the early detection and accurate diagnosis of breast cancer patients is significant to improve their prognosis. In this study, we propose a multi-modal, data-driven intelligent framework especially designed to breast cancer analysis the Multi-Dimensional Feature Refinement Cancer Network (MDFR-Cancer-Net) that can be used to analyse breast cancer at various clinical stages based on multi-modal data.The framework of this study will be to combine tabular clinical data for risk stratification, pathological imaging data of tumors and mi-RNA molecular data for biomarker-assisted analysis at different times in the study. MDFR-Cancer-Net will first be used to reduce the number of features and improve the representation ability of features. A few are Naive Bayes and Deep U-Net; the rest are other types of models for classification and segmentation. Five-fold cross-validation and ablation tests were carried out to obtain the above results. Based on the above experimental results, the accuracy of the Naive Bayes model on the Wisconsin Breast Cancer dataset was 99%, and that of the baseline SVM was 92%. Deep U-Net reached an accuracy of 96% for the 277,524 IDC pathological image patches and exceeded the baseline CNN's accuracy of 87%. Naive Bayes had 100% accuracy in the mi-RNA molecular analysis and outperformed the baseline Random Forest (95%). In short, the framework presented here shows that all kinds of data can be used together to help diagnose breast cancer.
Yun-Ze Li, N. Sani· Theoretical and Natural Scie...· 0 citations
A clear, statistically sound, yet easily understandable breast cancer diagnosis is a difficult issue in all healthcare systems, because early stages of breast cancer are critical in therapy success and long-term survivability. This machine-learning-based breast cancer classifier, in a statistically justified, rigorously experimentally validated way, classifies a set of 569 breast cancer cases with 9 cytological features for breast cancer diagnosis. The classifier uses a rigorous set of data cleanup measures, including missing-value substitution, correlation-based feature reduction, and projection into principal component space, to achieve high data quality, reduce redundancy, and enhance feature usefulness. Five supervised classifiers, in a widely accepted train-test model using an 80:20 random sample split and 5-fold cross-validation, are fitted and evaluated using Accuracy, Precision, Recall, F1-score, and Area under the ROC curve. In these tests, the Random Forest classifier got the best result, with 95.84% Accuracy, 95.31% Precision, 95.12% Recall, 95.21% F1-score and 0.982 area under the ROC curve; in a statistically sound consistency test using cross-validation, its mean accuracy reached 95.96% with a small standard deviation of 0.43. To provide a clear, interpretable indication of which features truly matter, we performed a feature-importance analysis on the best classifier, the Random Forest model. Results show that the expression levels of Bland Chromatin, Single Epithelial Cell Size, Normal Nucleoli, Uniformity of Cell Shape, Uniformity of Cell Size and Bare Nuclei are closely related to breast cancer diagnosis; this is almost the same as the clinical diagnosis findings, and very naturally suggests that abnormalities of cellular morphology and nuclei are major symptoms of breast cancer. In comparison, prior research may neglect validation and efficiency comparisons or focus only on the classifier's accuracy. Our method combines multiple levels of assessment (statistical data-by-data validation, feature importance, cross-validation, and comparison of different classifiers using ensemble learning) into a single evaluation system. This combined approach not only enhances predictive capability but also makes the entire setup more explicitly interpretable from a clinical perspective, thereby making it more suitable for health care decision support. Given the strong classification performance, interpretability, and validation suggested above, the model would help physicians detect breast cancer very early, reducing the risk of misdiagnosis.
T. Haripriya, M. V. Ramana Murthy, Ch. Vasavi et al.· International Journal of Eng...· 0 citations
One of the main causes of cancer-related fatalities globally is still lung cancer, and increasing survival rates depends on early identification. In this work, a machine learning-based method for predicting lung cancer utilising clinical and lifestyle data from surveys is presented. The Synthetic Minority Over-sampling Technique (SMOTE) was used to address the dataset's class imbalance and guarantee equitable representation of both classes. Logistic regression was used as the meta-learner in a stacked ensemble model that combined CatBoost, XGBoost, LightGBM, AdaBoost, and Random Forest. The suggested methodology outperformed individual classifiers and came close to state-of-the-art performance documented in the literature, with an accuracy of 96.9% and a ROC AUC of 0.99. A Flask-based web application that offers an intuitive interface for prediction, visualisation, and outcome interpretation was also developed. Modules for user registration, input-based cancer risk prediction, and graphical display of test data results are all included in the system. The findings show that ensemble learning provides a reliable and user-friendly method for lung cancer prediction when combined with efficient preprocessing and web deployment.
Bushra Khanam, Lubna Nausheen, F. Fatima· International Journal of Dat...· 0 citations
The results demonstrate that ensemble learning structures, which combine the strengths of multiple models, can provide more reliable decision support in critical areas such as healthcare and emphasise that ensemble learning structures, which combine the strengths of multiple models, can provide more reliable decision support in critical areas such as healthcare.
Yasin Karakuş, Pınar Özen· Journal of Innovative Engine...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.