An architecture-agnostic framework is proposed that augments VLM inputs with spatially localized, disc-level anomaly heatmaps generated by a semi-supervised U-Net++ model that improves anatomical sensitivity through explicit visual grounding and provides an independent interpretability output for clinical oversight, moving us closer to diagnostically reliable, visually grounded VLMs for lumbar spine MRI interpretation.
Abstract
Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for Vision-Language Models (VLMs). We benchmark state-of-the-art VLMs on lumbar spine MRI with a focus on diagnostic accuracy and demonstrate that standard lexical and semantic metrics poorly reflect clinical correctness: fluent, well-structured reports can score highly while containing clinically meaningful diagnostic errors. To address this failure mode, we propose an architecture-agnostic framework that augments VLM inputs with spatially localized, disc-level anomaly heatmaps generated by a semi-supervised U-Net++ model. These heatmaps both improve anatomical sensitivity through explicit visual grounding and provide an independent interpretability output for clinical oversight, moving us closer to diagnostically reliable, visually grounded VLMs for lumbar spine MRI interpretation.
This work presents Stroke CT Analysis and Natural Language Reporting (SCAN-R), a unified end-to-end framework that integrates multiclass stroke detection, Transformerenhanced U-Net segmentation with task-specific pre-trained backbones, and Retrieval-Augmented Generation for evidence-based clinical report generation.
Le Minh Toan Truong, X. Nguyen, Dang Khanh Tran· International Conference on...· 0 citations
In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that underpins response assessment, recurrence detection, and ongoing patient management. Yet,...
Kegeng Tang, Jing-Bo Wang, Shaogang Ren et al.· 0 citations
Automated radiology report generation (ARRG) has emerged as a promising application of artificial intelligence for reducing radiologists’ documentation workload and improving the consistency of clinical reporting. However, conventional image-to-text models often struggle to capture subtle abnormalities, establish meani...
P. Dayaker, M. Vignesh, I. Z. et al.· International journal of com...· 0 citations
Neuroradiologists rarely read a brain MRI in isolation, yet automated brain-MRI report generation has been built almost entirely for single studies. Temporal analysis has been explored on chest radiography and chest CT, but to our knowledge, longitudinal reporting for brain MRI, where interval change is often subtle an...
Magnetic Resonance Imaging (MRI) interpretation is fundamental to clinical decision-making, requiring radiologists to integrate multi-view anatomical planes across sequential timepoints while precisely localizing interval changes. However, existing vision-language benchmarks remain confined to single-timepoint, single-...
Wafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura et al.· 0 citations
Positron emission tomography-computed tomography (PET-CT) reporting is cognitively demanding, requiring integration of quantitative metabolic data and anatomical findings across multiple body regions. Existing segmentation and analysis models provide anatomical and physiological parameters, but not cohesive physici...
D. Prakash, Pramukh Kulkarni, Venkatesh Rangarajan et al.· Indian Journal of Nuclear Me...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.