Skip to content
Open access

Tree-Based Machine Learning for Diagnostic Classification of Dengue Fever Using Routine Hematological Parameters: A Secondary Analysis of a Publicly Available Dataset

Sep 2026 · Diagnostics · Vol 16 · 0 citations · 59 references
Medicine

TL;DR

Tree-based machine-learning models for dengue classification showed moderate discrimination, with high sensitivity but limited specificity, and yielded numerically higher AUROCs and lower Brier scores than the primary SMOTE-trained LR within this internal-validation framework.

Abstract

Background: Dengue fever remains a major global health problem, and early diagnosis is challenging where confirmatory testing is limited. Machine-learning studies using routine hematological data have focused mainly on discrimination, whereas calibration, decision-analytic performance, interpretability, and robust validation have received less attention. This study aimed to develop and compare tree-based machine-learning models for dengue classification, benchmark them against L2-penalized logistic regression (LR), and evaluate discrimination, calibration, potential decision-analytic benefit, and interpretability. Methods: This retrospective secondary analysis used an open-access dataset from Bangladesh comprising 1523 patients, 18 demographic and hematological predictors, and a binary dengue test outcome. Data were divided into stratified training (80%) and test (20%) sets. The Synthetic Minority Over-sampling Technique was applied only within the training workflow. Random Forest (RF), XGBoost, and LightGBM were optimized using Optuna with stratified five-fold cross-validation. L2-penalized LR was evaluated using the same predictors and training–test partition. Held-out test-set performance was assessed using AUROC, AUPRC, accuracy, sensitivity, specificity, predictive values, F1-score, and Brier score. Calibration, decision curve analysis, SHAP values, and permutation importance were also examined. Results: LightGBM, RF, and XGBoost yielded AUROCs of 0.709, 0.704, and 0.702, respectively, indicating closely similar discrimination. The primary SMOTE-trained LR yielded a numerically lower AUROC of 0.608 (95% CI: 0.536–0.677) and a higher Brier score of 0.310 than the tree-based models (0.174–0.176); however, in sensitivity analysis without SMOTE, the LR AUROC increased numerically to 0.655 and the Brier score decreased to 0.194. At the training-derived threshold of 0.558, LightGBM achieved a sensitivity of 0.914 and a specificity of 0.458, reflecting a high-sensitivity, low-specificity profile. The LightGBM calibration curve suggested closer agreement in the low-to-moderate predicted-probability range, with greater deviation at higher probabilities. Decision curve analysis suggested potential net benefit across a range of threshold probabilities but did not establish clinical utility. Platelet count, monocyte percentage, and neutrophil percentage were consistently among the leading predictors across the tree-based models. Conclusions: Tree-based models showed moderate discrimination, with high sensitivity but limited specificity, and yielded numerically higher AUROCs and lower Brier scores than the primary SMOTE-trained LR within this internal-validation framework. They should not replace etiological testing or be used as standalone diagnostic tools; their observed operating characteristics are more compatible with a potential adjunctive screening or triage-support role. External and prospective validation across independent populations and settings is required before clinical use or superiority over simpler statistical models can be established.

Read PDF

Similar papers

Sep 2026

Comparison of Tree-Based Machine Learning Models for Classification of Tuberculosis Outcomes in Brazil

Tuberculosis remains a significant public health concern, recognized as a reemerging disease strongly associated with socioeconomic conditions. According to the World Health Organization, tuberculosis continues to be the leading cause of death from a single infectious agent worldwide in 2025. This study evaluates Rando...

Heloísa de Almeida Pereira, Marcos Roberto Ribeiro, Ciniro Aparecido Leite Nametala · 0 citations
Open access Sep 2026

Predicting Dengue Clinical Severity in Eastern Sudan's 2023 Outbreak: A Comparative Analysis of Statistical and Machine Learning Models Using Routine Surveillance Data

The study-specific composite clinical severity indicator was uncommon but was associated with a higher risk of death, and the findings require confirmation using independent datasets with more detailed clinical and laboratory information.

Fathelrhman el Guma, Elkhatim Abuelysar, EihabAbdelhai Osman et al. · 0 citations
Open access Sep 2026

Pathogen prevalence, feature composition and cross-centre generalisability of machine learning diagnostic models for multi-pathogen respiratory infection

To evaluate the predictive information contained in the restricted surveillance feature set (demographic, temporal and specimen variables) for respiratory pathogen identification, and to examine factors associated with model performance. We retrospectively analysed 24,689 acute respiratory infection cases fr...

Feng-Miao Hu, Xing-Yu Zhou, Li-Jun Zhou et al. · 0 citations
Open access Aug 2026

An Explainable Transformer-Based Machine Learning Framework for Bilirubin-Derived Severity Classification of Hepatitis B: A Cross-Domain Validation Study

Background: Accurate and comprehensible severity evaluation is necessary for chronic hepatitis B (HBV), which continues to be a significant worldwide health concern. Complex nonlinear interactions among clinical biomarkers are frequently missed by conventional clinical ratings. Objective: Using a novel TabTransformer...

Surojit Sadhu, M. Maindarkar, M. Karthikeyan et al. · 0 citations
Open access Sep 2026

Machine learning models for predicting liver cancer: a real-world cohort study in China

The findings support the feasibility of leveraging large-scale inpatient laboratory data for risk-stratification model development and show good discrimination and interpretability for liver cancer prediction in a hospitalized real-world cohort using routinely available clinical and laboratory data.

Ce-Xiong Fu, Fang Li, Shi-Bing Li et al. · 0 citations
Review Open access Sep 2026

Early malaria risk screening in Nigerian minors using AutoML and cluster-based analysis of non-clinical survey data

Malaria remains a major public health challenge in sub-Saharan Africa, with Nigeria accounting for a substantial proportion of the global malaria burden, particularly among children under five years of age. Although rapid diagnostic tests (RDTs) enable timely screening, false-negative results and delays in confirmatory...

A. Khan, Nafisa Mahbub, Ridwan Al Aziz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.