A Compression-First Knowledge-Distilled Neural Network for ICU Stroke Mortality Prediction
Abstract
Accurate and timely prediction of in-Intensive-Care-Unit (ICU) mortality in stroke patients is essential for triage and resource allocation, yet the most accurate models reported on the Medical Information Mart for Intensive Care IV (MIMIC-IV) database are large ensembles or gradient-boosted trees whose memory footprint and inference cost hinder deployment at the bedside. In this work, a compression-first knowledge distillation (KD) framework is proposed in which an ensemble of three Transformer teachers transfers dark knowledge to a 1,633-parameter multilayer-perceptron student through a temperaturecorrected logit-matching objective augmented with FitNets-style feature hints. A stroke cohort of 12,408 ICU admissions was extracted from MIMIC-IV and processed with leakage-controlled Multiple Imputation by Chained Equations and patient-level splitting. On a held-out test set of 1,862 patients, the distilled student attains an AUROC of 0.931 (95% CI [0.914, 0.948]), which is statistically indistinguishable from the Transformer ensemble, XGBoost, LightGBM and Random Forest under the DeLong test ($p>0.28$ in all cases), while being $171 \times$ smaller than a single teacher and $2,939 \times$ smaller and $876 \times$ faster than the strongest tree baseline. A multi-seed capacity sweep and a 50-configuration hyperparameter analysis are reported to delineate the regime in which distillation contributes measurable value. The results demonstrate that bedside-deployable mortality scoring is achievable without sacrificing discriminative performance, and that rigorous statistical validation is indispensable when reporting marginal differences on clinical prediction tasks.