Skip to content
Preprint

Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report Generation

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

An architecture-agnostic framework is proposed that augments VLM inputs with spatially localized, disc-level anomaly heatmaps generated by a semi-supervised U-Net++ model that improves anatomical sensitivity through explicit visual grounding and provides an independent interpretability output for clinical oversight, moving us closer to diagnostically reliable, visually grounded VLMs for lumbar spine MRI interpretation.

Abstract

Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for Vision-Language Models (VLMs). We benchmark state-of-the-art VLMs on lumbar spine MRI with a focus on diagnostic accuracy and demonstrate that standard lexical and semantic metrics poorly reflect clinical correctness: fluent, well-structured reports can score highly while containing clinically meaningful diagnostic errors. To address this failure mode, we propose an architecture-agnostic framework that augments VLM inputs with spatially localized, disc-level anomaly heatmaps generated by a semi-supervised U-Net++ model. These heatmaps both improve anatomical sensitivity through explicit visual grounding and provide an independent interpretability output for clinical oversight, moving us closer to diagnostically reliable, visually grounded VLMs for lumbar spine MRI interpretation.

View source

Similar papers

Conference Aug 2026

SCAN-R: bridging precision imaging and natural language for automated stroke diagnosis

This work presents Stroke CT Analysis and Natural Language Reporting (SCAN-R), a unified end-to-end framework that integrates multiclass stroke detection, Transformerenhanced U-Net segmentation with task-specific pre-trained backbones, and Retrieval-Augmented Generation for evidence-based clinical report generation.

Le Minh Toan Truong, X. Nguyen, Dang Khanh Tran · 0 citations
#computer vision Preprint Aug 2026

CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that underpins response assessment, recurrence detection, and ongoing patient management. Yet,...

Kegeng Tang, Jing-Bo Wang, Shaogang Ren et al. · 0 citations
Open access Aug 2026

Attention-Guided Vision-Language Model for Automated Radiology Report Generation

Automated radiology report generation (ARRG) has emerged as a promising application of artificial intelligence for reducing radiologists’ documentation workload and improving the consistency of clinical reporting. However, conventional image-to-text models often struggle to capture subtle abnormalities, establish meani...

P. Dayaker, M. Vignesh, I. Z. et al. · 0 citations
Preprint Sep 2026

BrainDiff: Longitudinal Report Generation for Multimodal Brain MRI

Neuroradiologists rarely read a brain MRI in isolation, yet automated brain-MRI report generation has been built almost entirely for single studies. Temporal analysis has been explored on chest radiography and chest CT, but to our knowledge, longitudinal reporting for brain MRI, where interval change is often subtle an...

Krish Patel, Peirong Liu · 0 citations
Preprint Aug 2026

How Good are Foundation Models in Longitudinal MRI Disease Progression Reasoning?

Magnetic Resonance Imaging (MRI) interpretation is fundamental to clinical decision-making, requiring radiologists to integrate multi-view anatomical planes across sequential timepoints while precisely localizing interval changes. However, existing vision-language benchmarks remain confined to single-timepoint, single-...

Wafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura et al. · 0 citations
Review Open access Sep 2026

Context-Grounded PET-CT Report Generation Using LoRA-Fine-Tuned BioMedLM with Physician Validation and Safety Evaluation

Positron emission tomography-computed tomography (PET-CT) reporting is cognitively demanding, requiring integration of quantitative metabolic data and anatomical findings across multiple body regions. Existing segmentation and analysis models provide anatomical and physiological parameters, but not cohesive physici...

D. Prakash, Pramukh Kulkarni, Venkatesh Rangarajan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.