Skip to content
Review

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

Aug 2026 · International Journal of Medical Informatics · Vol 221, pp. 106684 · 0 citations · 24 references
Medicine

TL;DR

Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias.

Abstract

Purpose

Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias.

Methods

Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative.

Results

Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain.

Conclusion

Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

View source

Similar papers

Review Sep 2026

Mixture-of-experts and gating mechanisms in multimodal biomedical learning: a routing-centric analysis

A novel three-axis taxonomy covering Gating Architecture, Application Level, and Supporting Techniques is proposed and a controlled comparison of six gating architectures on three biomedical benchmarks shows that the five Axis 1 gating types produce measurably distinct performance profiles and that no single architectu...

Ba-Duy Nguyen, Van-Dung Hoang, Hien D. Nguyen · 0 citations
Review Open access Aug 2026

Deep learning and generative AI for medical imaging and clinical decision support systems: a structured critical review

It is argued that LLM-based CDSS are not supported for routine autonomous use and require clinician supervision, and set out a research agenda centred on validation, governance and human-in-the-loop deployment.

A. Babu, A. J. Nehemiah, V. Jagadeep et al. · 0 citations
Open access Aug 2026

CARE-LLM-GRAPH: Confidence Aware LLM integrated Multimodal Architecture for clinical Recommendation

A new confidence-aware hybrid design, CARE-LLM-GRAPH, which combines large language models (LLMs) to perform clinical reasoning, multimodal deep learning to analyze medical images, and population-aware graph intelligence to provide cohort-level information is presented.

Unknown authors · 0 citations
Review Aug 2026

Hybrid deep learning for breast cancer diagnosis: a mechanism-centered systematic review, taxonomy, and evidence synthesis

This systematic review synthesizes hybrid deep learning methods across mammography, ultrasound, MRI, histopathology/whole-slide imaging, and clinically oriented multimodal settings and provides practical reporting and evaluation recommendations to support more reliable and clinically ready systems.

Shweta Sharma, A. Jatain, J. Kataria et al. · 0 citations
#large language models Review Open access Sep 2026

Multimodal medical diagnosis: a mini review of LLM–vision fusion models in low-resource healthcare settings

Recent advances in large language models (LLMs) and vision transformers have enabled multimodal systems that integrate clinical text with medical imaging for diagnostic decision-making. While these systems show promising results on benchmark datasets in well-resourced research settings, their applicability in low-resou...

Kahakashan Ashraf, Md.Hamid Hosen, N. Farah et al. · 0 citations
Conference Open access Sep 2026

Towards Reliable Multimodal Clinical Decision Support: From Quality-Aware Fusion to Patient Digital Twins

This doctoral research addresses fragility in multimodal deep learning by shifting the paradigm from generative imputation to uncertainty-guided abstention and introducing QA-MoE (Quality-Aware MoE), which explicitly decouples reliability estimation from the routing process.

Lin-Peng Sun · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.