Skip to content
Review

Large language models for lumbar spondylolisthesis detection: a multi-center pilot comparative radiographic accuracy study

Aug 2026 · European Journal of Orthopaedic Surgery & Traumatology · Vol 36 · 0 citations · 36 references
Medicine

TL;DR

ChatGPT-5 outperformed ChatGPT-4o in detecting lumbar spondylolisthesis and both LLMs remained limited compared to fellowship trained spine surgeons.

View source

Similar papers

Open access Aug 2026

Prompt Configurations for Multimodal Large Language Models in Diagnosing and Staging Osteonecrosis of the Femoral Head: Multimodel Retrospective Observational Diagnostic Study

The findings support the potential role of MLLMs as assistive tools in human-AI orthopedic imaging workflows, but external validation, careful input standardization, and prospective clinical evaluation are needed before clinical deployment.

Jiesheng Zhu, Xing-Xing Huang, Jin-Cheng Shi et al. · 0 citations
Open access Aug 2026

Multimodal Large Language Models vs. Medical Doctors in Degenerative Lumbar Spine Surgery: A Retrospective Decision Concordance Study of 147 Patients

Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization, establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.

M. Hamdan, A. Harati, A. Al-bakheet et al. · 0 citations
Review Open access Sep 2026

Diagnostic performance of GPT-5.2 ınstant for pediatric elbow fracture detection on radiographs: a prospective single-center diagnostic accuracy study

Multimodal large language models can interpret medical images, but their performance for pediatric elbow radiographs remains uncertain. We evaluated the diagnostic performance of GPT-5.2 Instant as accessed through the ChatGPT web interface during the defined study period. In this prospective, single-cente...

O. Taş, Mehmet Yorgun, R. Aktaş et al. · 0 citations
Open access Sep 2026

Large Language Models for Ankle Fracture Classification and Management Prediction from Routine Clinical Documentation: A Single-Center Exploratory Study.

It is suggested that LLMs can extract structured information from routine clinical documentation, performing well for standardized classification but less reliably for procedure-level prediction.

Benjamin Schwarberg, C. Ketzer, B. Thiel et al. · 0 citations
Open access Aug 2026

Multicenter evaluation of four large language models for automated spine imaging diagnosis

Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Her...

Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.