Skip to content
Open access

Causal-Pathway-Guided DNN–GBDT Distillation for Interpretable Artificial Intelligence in Intensive Care Units

Jul 2026 · Informatics · Vol 13, pp. 124 · 0 citations · 40 references

TL;DR

The proposed causal-aware distilled GBDT achieves stronger predictive performance than conventional interpretable baselines and substantially higher causal consistency than black-box temporal models, suggesting that causal structure can serve as an inductive bias for converting complex temporal prediction into interpretable rule-based clinical reasoning.

Abstract

Artificial intelligence (AI) systems for intensive care units (ICUs) must support early risk prediction while producing explanations that clinicians can inspect, question, and relate to physiological reasoning. Deep neural networks (DNNs) can learn complex temporal patterns from electronic health records (EHRs), but their internal representations are often difficult to translate into clinically actionable explanations. Gradient-boosted decision trees (GBDTs) offer more transparent decision rules, yet they may not capture the full temporal and nonlinear structure of high-dimensional ICU data. This paper presents a causal-pathway-guided DNN–GBDT distillation framework for interpretable ICU decision support. The framework first estimates a directed acyclic graph (DAG), denoted by G, from multivariate ICU time-series data and then uses the graph to guide representation learning in a DNN teacher model through causal gating. The learned teacher is distilled into a GBDT student model using soft predictive targets and a causal attribution-guided split-selection procedure, so that the final model approximates the teacher predictions while prioritizing tree splits aligned with plausible physiological pathways. Experiments using Medical Information Mart for Intensive Care IV (MIMIC-IV) data evaluate sepsis onset and in-hospital mortality prediction through discrimination, precision–recall performance, calibration-oriented reporting, causal consistency, and clinical utility indicators. The proposed causal-aware distilled GBDT achieves stronger predictive performance than conventional interpretable baselines and substantially higher causal consistency than black-box temporal models. The results suggest that causal structure can serve as an inductive bias for converting complex temporal prediction into interpretable rule-based clinical reasoning. The paper also discusses limitations related to observational causal discovery, unmeasured confounding, temporal stationarity, and clinical deployment, following recent reporting expectations for AI-based clinical prediction models.

Read PDF

Similar papers

Aug 2026

GuardMLLM: Overconfidence-Aware Dynamic Fusion for Early Outcome Prediction.

Experimental results show that the proposed GuardMLLM improves performance on tasks such as predicting patient mortality and ICU length of stay, and effectively alleviates overconfidence in LLM.

Gang Fu, Xiao-Long Xu, Hao-Long Xiang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

INTERVenE: Temporal-Abstraction-Interval Based Transformers for Short-Horizon Medical Event Prediction

INTERVenE is presented, a family of Transformer architectures whose input is an interval-based, knowledge-based temporal abstraction (KBTA), a token stream of named clinical concepts drawn from a curated medical ontology, rather than an unnamed bin index or a raw measurement triplet.

Shahar Oded, Yuval Shahar · 0 citations
Open access Sep 2026

Explainable Hybrid GRU–TabTransformer Learning with Cross-Attention and LLM-Assisted Interpretation for Stroke Risk Prediction

Early stroke risk prediction offers an opportunity for timely interventions and may help reduce the clinical burden associated with stroke. Artificial intelligence (AI) provides medical practitioners with tools to analyse clinical biomarkers and predict a patient’s stroke risk. However, existing models lack interpretab...

Moses Guddah, Adham Atyabi · 0 citations
Preprint Aug 2026

Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records

BERT-LER is presented, a BERT-style model for coded EHR timelines pretrained and fine-tuned from a de-identified EHR dataset of 75 million patients, that encodes laboratory test results as discrete tokens while retaining graded information through percentile-based binning, paired with Integrated Gradients for token-lev...

Jun-Ni Du, Lukas Adamek, Maxim A Kryukov et al. · 0 citations
Conference Aug 2026

A Compression-First Knowledge-Distilled Neural Network for ICU Stroke Mortality Prediction

Accurate and timely prediction of in-Intensive-Care-Unit (ICU) mortality in stroke patients is essential for triage and resource allocation, yet the most accurate models reported on the Medical Information Mart for Intensive Care IV (MIMIC-IV) database are large ensembles or gradient-boosted trees whose memory footprin...

Habibur Rahaman, Mohammad Shamsul Arefin · 0 citations
Open access Jul 2026

A Neuro-Symbolic Knowledge Graph and Large Language Model Hybrid Architecture for Multi-Modality Mental Health Counseling

A neuro-symbolic architecture achieves near-perfect guideline-appropriate routing with a governance profile, traceability, reproducibility, machine-traceable provenance, externally validated crisis detection, and deterministic provider control aligned with requirements for regulated clinical AI.

J. Tao, N. Fenn, H. Parent et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.