Skip to content
Review Open access

Multimodal Large Language Models for Prognostic Prediction in Cervical Cancer Treated with Definitive Chemoradiotherapy: An Exploratory Study of Systematic Multimodal Data Integration

Sep 2026 · Cancers · 0 citations · 33 references

Abstract

Background/Objectives: About 25–35% of patients with cervical cancer treated with definitive concurrent chemoradiotherapy (CCRT) relapse within five years, and non-imaging clinical factors stratify their risk only modestly. We evaluated whether a general-purpose multimodal large language model (MLLM), used without fine-tuning, could estimate recurrence risk in this setting. Methods: In this retrospective single-center study, 82 patients treated with definitive radiotherapy (79/82 with concurrent platinum-based chemotherapy) who had a complete pretreatment MRI report were analyzed. Gemini 3.1 Pro generated a structured report from pelvic MRI images and, separately, estimated recurrence/metastasis risk from clinical, laboratory, treatment and imaging data. Because treatment cycles and overall treatment time actually completed were included, this was a retrospective treatment-complete risk assessment rather than a strictly pretreatment prediction. Five prompting strategies differing only in their inputs—an ablation of input modalities—were each run three times; AI reports were graded against paired human reports on a 14-item rubric by an independent model (Claude Opus 4.6). Results: Forty patients (48.8%) relapsed. Discrimination rose from an AUC of 0.704 with a clinical baseline to 0.790 with the full multimodal input (95% CI 0.687–0.883; ΔAUC +0.085). This gain was significant on our primary two-sided bootstrap test after Holm correction for ten pairwise comparisons (p = 0.032; a DeLong sensitivity analysis supported some but not all the imaging-benefit comparisons). Adding MRI report text significantly improved the AUC over the clinical baseline, and no statistically significant difference was detected between the human-report and AI-report strategies; this was not an equivalence or non-inferiority test; and the further lymph-node increment was not statistically significant. The AI reports scored 68.8% against an LLM judge on the rubric, with apparent overcalling of parametrial invasion and an inability to assess lymph nodes within the supplied field of view; all strategies underestimated absolute risk (E/O 0.78–0.83). Conclusions: A non-fine-tuned MLLM can integrate multimodal data into a prognostic estimate, but its automatically generated MRI reports are frequently discordant with radiologist reports in specific ways and require expert review and external validation before clinical use.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.