A systematic review of artificial intelligence and machine learning for public health predictions using electronic health records in low and middle income countries with comparative benchmarks from high income settings
The growing availability of electronic health records (EHRs) has accelerated the use of artificial intelligence (AI) and machine learning (ML) in public health. Yet, how well these methods work in low- and middle-income countries (LMICs), remains poorly understood. This review synthesises studies on ML-based prediction using EHR or routinely collected electronic clinical data, with direct LMIC evidence analysed alongside high-income country (HIC) studies, included as methodological benchmarks. Following PRISMA guidelines, searches across five major databases identified 64 eligible studies published between Jan 2018–Mar 2025. Of these, 12 (18.8%) were conducted exclusively in LMIC settings, 44 (68.8%) in HICs, and 8 (12.5%) drew on mixed or multi-setting data. Retrospective designs predominated (81.3%). Disease progression (40.6%), mortality (34.4%), and treatment response (25.0%) were the most common prediction targets. Deep learning architecture was the most frequently applied category overall (39.1%, n = 25), driven by HIC studies with access to large curated datasets; among LMIC-focused studies, traditional ML and ensemble methods were each applied in 33.3% of studies. Evaluation practices were dominated by discrimination metrics, particularly AUROC; external validation was reported in only 5 studies (7.8%) and calibration in only 4 (6.2%). Explainability assessment was reported in 1 of 12 LMIC studies (8.3%) compared with 16 of 44 HIC studies (36.4%), with governance and ethical considerations inconsistently documented in LMIC settings. This review highlights key methodological and contextual gaps and offers guidance for developing interpretable, reliable, and context-appropriate AI tools for public health decision-making in LMIC settings.