Examination of Diabetes Prediction Using Machine Learning
Abstract
: Diabetes has become a severe challenge to global public health: in 2023, there are 537 million cases worldwide (expected to rise to 783 million by 2045), with annual medical expenses exceeding $727 billion. Furthermore, 30% to 50% of patients are undiagnosed and asymptomatic, highlighting the urgent need for precise predictive models. The advantages of machine learning are significant: for instance, the random forest model combining HbA1c and FLI achieved an AUC of 0.874, while the cross-population model from the THIN database achieved an AUC ranging from 0.907 to 0.925. The integration of lifestyle data through CATBoost revealed a U-shaped risk association with sleep duration, and the accuracy of the multi-source hybrid model exceeded 98%, confirming the value of multimodal integration. However, current research faces challenges including data heterogeneity, insufficient external validation, poor interpretability of deep learning, and difficulties in multimodal integration. In the future, it is essential to establish a standardized validation framework, develop interpretable algorithms, integrate wearable non-invasive markers, and implement Bayesian racial modeling to promote early screening and personalized intervention, thereby revolutionizing the clinical prevention paradigm.