MACHINE LEARNING IN PREDICTION OF SEPSIS IN INTENSIVE CARE UNITS: A LITERATURE REVIEW OF ALGORITHMIC MODELS AND CLINICAL IMPLEMENTATION
Abstract
Background. Sepsis continues to be among the leading causes of in-hospital death, accounting for around 11 million deaths globally in 2017 and pooled ICU mortality of nearly 42% (Rudd et al., 2020; Fleischmann-Struzek et al., 2020). The chance of survival decreases with each additional hour of delay in starting effective antibiotic treatment (Kumar et al., 2006), and standard screening tools, including SIRS and qSOFA, are known to overlook a meaningful share of cases (Raith et al., 2017). Models trained on routinely captured EHR data have therefore been put forward as a means of recognising patients at risk well before overt deterioration. Objective. To bring together the peer-reviewed evidence published between 2015 and 2025 on machine-learning algorithms for predicting sepsis in adult ICU patients, with attention to model architecture, prospective validation and the regulatory status of systems that have actually reached the bedside. Methods. A PRISMA 2020-aligned narrative review (Page et al., 2021) covering MEDLINE, Scopus, Web of Science, Embase and IEEE Xplore from January 2015 to March 2025. Reporting quality was checked against TRIPOD+AI (Collins et al., 2024) and risk of bias with PROBAST (Wolff et al., 2019). Seventy-three studies met the inclusion criteria. Results. Gradient-boosted trees, LSTM networks and transformer-based models reach internal AUROCs of about 0.83 to 0.94, while qSOFA and SIRS sit between 0.69 and 0.76 (Henry et al., 2015; Nemati et al., 2018). TREWS was associated with an 18.7% relative drop in adjusted in-hospital sepsis mortality (Adams et al., 2022); the only published randomised trial of an ML sepsis predictor, InSight, lowered mortality from 21.3% to 8.96% in a small ICU cohort (Shimabukuro et al., 2017); COMPOSER showed a 17% relative reduction at the bedside (Boussina et al., 2024). By contrast, the widely deployed Epic Sepsis Model achieved an external AUROC of only 0.63 (Wong et al., 2021). Two AI sepsis diagnostics have so far cleared the FDA: Sepsis ImmunoScore (De Novo DEN230036, 2024) and TriVerity (510(k) K241676, 2025). Conclusions. On the whole ML models discriminate sepsis better than traditional scores, but uneven outcome definitions, training-set bias, alert fatigue and patchy external validation still limit how far the gains travel between centres. Prospective trials reported to TRIPOD+AI standards, together with explicit equity audits, will be needed before ML sepsis prediction can be treated as routine care.