Aug 2026· The spine journal· 0 citations· 33 references
Medicine
TL;DR
OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT, although accuracy declined when outputs were reformatted through ChatGPT, suggesting cross-model contamination.
Abstract
Background
Context
Generative artificial intelligence (AI) is increasingly used in spine care; however, concerns remain regarding citation hallucinations and reliability. ChatGPT may generate inaccurate or fabricated references, whereas OpenEvidence (OE) prioritizes verified, peer-reviewed literature. This is the first study comparing OE and ChatGPT using cervical spine clinical guideline (CSCG) queries.
Purpose
To compare guideline alignment, citation validity, sourcing, and prompt-engineering effects between OE and ChatGPT using CSCGs.
STUDY DESIGN/
Setting
Cross-sectional comparative analysis.
PATIENT SAMPLE
No patient population was included.
OUTCOME MEASURES
Primary outcomes were guideline alignment score and citation validity (fully correct, partially hallucinated, or fully hallucinated). Secondary outcomes included source type, publication year, proportion published after CSCG release, and prompt-engineering effects.
Methods
A total of 110 evidence-based clinical questions derived from 10 CSCGs authored by 6 academic societies were submitted to OE (v2.0) and ChatGPT-4o from June 1 to July15, 2025. A subset of prompts was repeated to evaluate prompt-engineering effects.
Results
OE generated 999 citations with 100% accuracy, whereas only 184/393 (46.8%) ChatGPT citations met accuracy criteria (p<0.001). OE demonstrated higher guideline alignment than ChatGPT (4.6 ± 0.8 vs 4.1 ± 0.7; p = 0.03), with almost perfect interrater agreement (weighted Cohen's κ = 0.88; 95% CI, 0.80-0.97). . ChatGPT produced 88 partially hallucinated citations (22.4%), most commonly due to incorrect hyperlinks, author names, or publication years. OE cited more peer-reviewed literature (79.8% vs 60.1%; p<0.001) and more recent studies (2018±5.6 vs 2012±7.4; p<0.001). Prompt-engineering analysis showed OE maintained higher citation validity and fewer hallucinations, although accuracy declined when outputs were reformatted through ChatGPT, suggesting cross-model contamination.
Conclusion
OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT. Nonetheless, physician oversight remains essential for safe AI integration into clinical practice.
The findings describe informational quality, not clinical safety or decision-making, and support specialist verification before clinical use, and support specialist verification before clinical use.
Background: The reliability of generative artificial intelligence for Chinese medical information remains uncertain. This study evaluates the concordance and linguistic performance of three generative artificial intelligence systems in generating Chinese information on Ménière’s disease against clinical practice guidel...
While both models show potential as clinical decision-support tools, Gemini 3.0 exhibited superior diagnostic performance in real-world otologic cases in the first study evaluating LLMs using real-world otologic data.
Bilge Tuna, Gokhan Tuzemen, Hasan Mutlu· BMC Medical Informatics and...· 0 citations
ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery.
Kamil Balaban, Mehmet Batu Ertan, Mahmut Kalem· Digital Health· 0 citations
Objective
The aim of this study was to analyze the accuracy, reliability, and quality of the content of responses provided by artificial intelligence (AI)-based chatbots to frequently asked questions related to scoliosis surgery.
Methods
A set of 25 questions related to diagnosis, treatment options, surgical risks, a...
G. Alibakan, Y. Sulek· Cirugía y Cirujanos· 0 citations
ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable, but the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability obse...
Furkan Türkoğlu, Elif Nur Gencer, Emre Erdoğan· Archives of Current Medical...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.