Skip to content
Review

Evolution of Generative Artificial Intelligence in Clinical Practice: Comparative Performance of OpenEvidence 2.0 and ChatGPT-4o.

Aug 2026 · The spine journal · 0 citations · 33 references
Medicine

TL;DR

OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT, although accuracy declined when outputs were reformatted through ChatGPT, suggesting cross-model contamination.

Abstract

Background

Context

Generative artificial intelligence (AI) is increasingly used in spine care; however, concerns remain regarding citation hallucinations and reliability. ChatGPT may generate inaccurate or fabricated references, whereas OpenEvidence (OE) prioritizes verified, peer-reviewed literature. This is the first study comparing OE and ChatGPT using cervical spine clinical guideline (CSCG) queries.

Purpose

To compare guideline alignment, citation validity, sourcing, and prompt-engineering effects between OE and ChatGPT using CSCGs. STUDY DESIGN/

Setting

Cross-sectional comparative analysis. PATIENT SAMPLE No patient population was included. OUTCOME MEASURES Primary outcomes were guideline alignment score and citation validity (fully correct, partially hallucinated, or fully hallucinated). Secondary outcomes included source type, publication year, proportion published after CSCG release, and prompt-engineering effects.

Methods

A total of 110 evidence-based clinical questions derived from 10 CSCGs authored by 6 academic societies were submitted to OE (v2.0) and ChatGPT-4o from June 1 to July15, 2025. A subset of prompts was repeated to evaluate prompt-engineering effects.

Results

OE generated 999 citations with 100% accuracy, whereas only 184/393 (46.8%) ChatGPT citations met accuracy criteria (p<0.001). OE demonstrated higher guideline alignment than ChatGPT (4.6 ± 0.8 vs 4.1 ± 0.7; p = 0.03), with almost perfect interrater agreement (weighted Cohen's κ = 0.88; 95% CI, 0.80-0.97). . ChatGPT produced 88 partially hallucinated citations (22.4%), most commonly due to incorrect hyperlinks, author names, or publication years. OE cited more peer-reviewed literature (79.8% vs 60.1%; p<0.001) and more recent studies (2018±5.6 vs 2012±7.4; p<0.001). Prompt-engineering analysis showed OE maintained higher citation validity and fewer hallucinations, although accuracy declined when outputs were reformatted through ChatGPT, suggesting cross-model contamination.

Conclusion

OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT. Nonetheless, physician oversight remains essential for safe AI integration into clinical practice.

View source

Similar papers

Sep 2026

Guideline-Based Evaluation of Five Generative Artificial Intelligence Chatbot Platforms for Clinician-Oriented Questions on Open Temporomandibular Joint Surgery.

The findings describe informational quality, not clinical safety or decision-making, and support specialist verification before clinical use, and support specialist verification before clinical use.

Selin Gaş, Erdinç Sulukan, Büşra Korkmaz · 0 citations
Open access Sep 2026

Multidimensional Evaluation of Guideline-Based Generative Artificial Intelligence Responses in Traditional Chinese: A Ménière’s Disease Study

Background: The reliability of generative artificial intelligence for Chinese medical information remains uncertain. This study evaluates the concordance and linguistic performance of three generative artificial intelligence systems in generating Chinese information on Ménière’s disease against clinical practice guidel...

Mien-Jen Lin, Yun-Chiao Wen, Li-Chun Hsieh et al. · 0 citations
Review Open access Aug 2026

Artificial intelligence in clinical decision-making: a comparison of ChatGPT 5.0 and Gemini 3.0 in otologic cases

While both models show potential as clinical decision-support tools, Gemini 3.0 exhibited superior diagnostic performance in real-world otologic cases in the first study evaluating LLMs using real-world otologic data.

Bilge Tuna, Gokhan Tuzemen, Hasan Mutlu · 0 citations
Open access Sep 2026

Accuracy and limitations of artificial intelligence chatbots in answering patient questions on scoliosis surgery.

Objective The aim of this study was to analyze the accuracy, reliability, and quality of the content of responses provided by artificial intelligence (AI)-based chatbots to frequently asked questions related to scoliosis surgery. Methods A set of 25 questions related to diagnosis, treatment options, surgical risks, a...

G. Alibakan, Y. Sulek · 0 citations
Open access Sep 2026

Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2

ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable, but the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability obse...

Furkan Türkoğlu, Elif Nur Gencer, Emre Erdoğan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.