Sep 2026· ACM Transactions on Computing for Healthcare· 0 citations· 48 references
TL;DR
This resource paper introduces SpineXR-VQA, an open-source, clinically grounded, and verified benchmark comprising 2,187 X-rays and 8,272 expert-verified, open-ended Question-Answer pairs that benchmark 15 state-of-the-art Multimodal Large Language Models (MLLMs), including proprietary systems such as Claude Sonnet, GPT-o4 mini, and Gemini Flash.
Abstract
Recent progress in Medical Visual Question Answering (VQA) has significantly aided clinical decision-making across various domains such as pneumonia, oncology, and neurology. However, spinal and musculoskeletal ailments remain critically underexplored. The development of reliable VQA models for spinal imaging is currently hindered by a lack of datasets and evaluation protocols that do not reflect the descriptive, diagnostic interpretations used in clinical practice. In this resource paper, we introduce SpineXR-VQA, an open-source, clinically grounded, and verified benchmark comprising 2,187 X-rays and 8,272 expert-verified, open-ended Question-Answer pairs. Unlike standard classification-based datasets, SpineXR-VQA features six expert-validated categories: abnormality, severity, location, diagnosis, treatment, and reasoning. Ten orthopedic specialists from India and Thailand validated all pairs, ensuring geographic diversity and high inter-rater agreement (Cohen's Kappa: 0.96 for questions, 0.93 for answers). We benchmark 15 state-of-the-art Multimodal Large Language Models (MLLMs), including proprietary systems such as Claude Sonnet, GPT-o4 mini, and Gemini Flash, alongside medical-specific and general open-weight models. While these models achieve moderate semantic alignment (median similarity: 0.72), a detailed analysis reveals that they consistently fail to capture essential diagnostic nuances, such as anatomical fidelity and clinical completeness. These results underscore the necessity for specialized models in spinal VQA, a gap SpineXR-VQA fills.
Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Her...
Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al.· npj Digital Medicine· 0 citations
Abstract Vision-language models (VLMs) are increasingly applied to medical imaging, yet public benchmarks may reward memorization over perception: their images and questions can enter pretraining corpora, and many items remain answerable from question text alone. We present an automated, agent-driven pipeline that buil...
Bo Liu, Han Gu, Xiang-Rui Li et al.· Research Square· 0 citations
Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignm...
Ling-Xuan Hou, Yu-Hua Xie, Yue Hu et al.· 0 citations
PURPOSE
Most clinical evaluations of large language models assess factual recall rather than the multi-step reasoning behind operative plans. Reasoning-tuned models, post-trained to generate explicit intermediate reasoning before answering, may better approximate surgical decision-making. We compared two such models fr...
C. Lam, Conor T. Boylan, Adit Ravishankar et al.· Spine Deformity· 0 citations
These findings support continued benchmarking and cautious, specialist-supervised use of multimodal LLMs in retinal education and structured imaging question answering and indicate that retinal image interpretation remains a major limitation.
Xiao-Chen Gu, Yi-Zhou Yang, Xuan-Qiao Lin et al.· Frontiers in Medicine· 0 citations
An architecture-agnostic framework is proposed that augments VLM inputs with spatially localized, disc-level anomaly heatmaps generated by a semi-supervised U-Net++ model that improves anatomical sensitivity through explicit visual grounding and provides an independent interpretability output for clinical oversight, mo...
Bruno Palau, Franziska Vogt, Daria Laslo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.