Aug 2026· International Journal of Medical Informatics· Vol 221, pp.
106684
· 0 citations· 24 references
Medicine
TL;DR
Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias.
Abstract
Purpose
Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias.
Methods
Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative.
Results
Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain.
Conclusion
Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.
A novel three-axis taxonomy covering Gating Architecture, Application Level, and Supporting Techniques is proposed and a controlled comparison of six gating architectures on three biomedical benchmarks shows that the five Axis 1 gating types produce measurably distinct performance profiles and that no single architectu...
Ba-Duy Nguyen, Van-Dung Hoang, Hien D. Nguyen· Network Modeling Analysis in...· 0 citations
It is argued that LLM-based CDSS are not supported for routine autonomous use and require clinician supervision, and set out a research agenda centred on validation, governance and human-in-the-loop deployment.
A. Babu, A. J. Nehemiah, V. Jagadeep et al.· Frontiers in Digital Health· 0 citations
A new confidence-aware hybrid design, CARE-LLM-GRAPH, which combines large language models (LLMs) to perform clinical reasoning, multimodal deep learning to analyze medical images, and population-aware graph intelligence to provide cohort-level information is presented.
Unknown authors· European Journal of Prosthod...· 0 citations
This systematic review synthesizes hybrid deep learning methods across mammography, ultrasound, MRI, histopathology/whole-slide imaging, and clinically oriented multimodal settings and provides practical reporting and evaluation recommendations to support more reliable and clinically ready systems.
Shweta Sharma, A. Jatain, J. Kataria et al.· Network Modeling Analysis in...· 0 citations
Recent advances in large language models (LLMs) and vision transformers have enabled multimodal systems that integrate clinical text with medical imaging for diagnostic decision-making. While these systems show promising results on benchmark datasets in well-resourced research settings, their applicability in low-resou...
Kahakashan Ashraf, Md.Hamid Hosen, N. Farah et al.· Frontiers in Digital Health· 0 citations
This doctoral research addresses fragility in multimodal deep learning by shifting the paradigm from generative imputation to uncertainty-guided abstention and introducing QA-MoE (Quality-Aware MoE), which explicitly decouples reliability estimation from the routing process.
Lin-Peng Sun· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.