Aug 2026· Skeletal Radiology· 0 citations· 20 references
Medicine
TL;DR
Three flagship OpenAI deployments achieved cartilage grading agreement indistinguishable from human inter-rater variability, while cost-optimized and open-weight variants performed measurably below.
Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Her...
Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al.· npj Digital Medicine· 0 citations
Background: Large language models (LLMs) have demonstrated strong performance on standardized medical examinations, with recent studies reporting performance approaching or exceeding that of senior medical residents. However, examination accuracy alone does not establish how models arrive at their answers or the relati...
F. Gafoor, M. Syed, M. Halai et al.· medRxiv· 0 citations
Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification, but none of the models matched senior radiologist-level performance.
Objective: To develop and evaluate an AI-assisted MRI method for quantitative knee cartilage morphometry in a multicenter phase III knee osteoarthritis trial. Methods: AI pre-segmentation used 3D full-resolution nnU-Net. Version 1.0 used separate femorotibial- and patellar-cartilage models, whereas version 2.0 used a u...
Bin-Bin Yang, Rui Huang, Yuan-Jing Xu et al.· 0 citations
This study provides a benchmark of LLM performance for radiology OSCE-style content generation and evaluation during a specific snapshot of artificial intelligence development (August 2024).
Ankush Ankush, Samriddhi Burman, Sydney Smith et al.· Radiology Advances· 0 citations
The findings support the potential role of MLLMs as assistive tools in human-AI orthopedic imaging workflows, but external validation, careful input standardization, and prospective clinical evaluation are needed before clinical deployment.
Jiesheng Zhu, Xing-Xing Huang, Jin-Cheng Shi et al.· Journal of Medical Internet...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.