Skip to content
Open access

Comparative Evaluation of Large Language Model Interfaces in Third Molar Surgery Complication Scenarios: Response Quality, Clinical Content, Potential Clinical Risk, Readability, and Externally Observable Response Latency—A Cross-Sectional Comparative Benchmark Study

Sep 2026 · Journal of Clinical Medicine · Vol 15, pp. 7397 · 0 citations · 27 references

TL;DR

In this single-generation benchmark, the sampled outputs showed different profiles in specialist-rated quality, clinical content, potential clinical risk, readability, and externally observable response latency under the specific interface configurations tested.

Abstract

Background/Objectives: Large language models (LLMs) are increasingly investigated in dentistry and oral and maxillofacial surgery, but comparative evidence regarding specialist-level response quality, clinical content, and potential safety concerns remains limited. This study compared four contemporary LLM interfaces in third molar surgery complication scenarios. Methods: Thirty open-ended specialist-level scenarios were submitted once to ChatGPT 5.4 Thinking, DeepSeek-V3.2, Gemini 3.1 Pro, and Microsoft Copilot (Work and Learning mode) in isolated single-turn sessions, yielding 120 responses. Two oral and maxillofacial surgeons, blinded to interface identity, independently assessed responses using a task-adapted five-point Global Quality Score (GQS). An additional post hoc criterion-referenced evaluation assessed diagnosis, management, escalation/referral, potentially harmful recommendations, safety-critical omissions, and potential clinical risk. Readability and externally observable response latency were also evaluated. Results: Among the sampled outputs, mean GQSs were 4.53 for Gemini, 4.40 for ChatGPT, 4.30 for DeepSeek, and 3.92 for Copilot (p < 0.001). Mean Clinical Content Scores were 4.80/5 for ChatGPT, 4.77 for Gemini, 4.75 for DeepSeek, and 3.90 for Copilot (p < 0.001). All responses were rated as recognizing the target complication. Within the 120 sampled outputs, one Copilot response met the study-specific criterion for a potentially harmful recommendation, and two Copilot responses met the predefined criteria for safety-critical omissions. Within this benchmark, the sampled Copilot outputs had higher study-specific potential clinical-risk scores than the sampled outputs from the other evaluated interfaces (p < 0.001). Readability and externally observable response latency also differed across the sampled interface outputs. Conclusions: In this single-generation benchmark, the sampled outputs showed different profiles in specialist-rated quality, clinical content, potential clinical risk, readability, and externally observable response latency under the specific interface configurations tested. These findings apply to the observed responses from the single synchronized testing session and should not be interpreted as establishing superiority of any underlying LLM, independent diagnostic accuracy, clinical equivalence, clinical safety, effectiveness, or readiness for autonomous decision support.

Read PDF

Similar papers

Open access Sep 2026

Evaluation of large language model responses to expert questions in anterior implant dentistry: quality, accuracy, and readability

AI chatbots can generate information with potential clinical relevance in anterior implant dentistry; however, variability in informational reliability persists and expert supervision remains essential before integrating such tools into clinical education.

Dalndushe Abdulai, Raghıb Suradı, Mehran Moghbel · 0 citations
Aug 2026

Comparative Evaluation of Large Language Models in Answering Patient Questions Following Periodontal and Peri-Implant Examination: An Expert-Based Study.

G Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations, which support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical...

Ramazan Ağırağaç, Vedat Yüksekkaya · 0 citations
Review Open access Oct 2026

Performance of five large language models in frozen shoulder patient education: a multidimensional evaluation of readability, quality, clinical alignment, and safety

This study compared the educational suitability, overall quality, translation-based readability, clinical intent alignment, and safety of frozen shoulder patient education materials generated by GPT-5, DeepSeek, Doubao, Wenxin Yiyan, and Tongyi Qianwen. Twenty clinician-curated questions covering five content categorie...

Jia-Chen Liang, Wan-Ying Gui, Hua-Nan Li · 0 citations
#small language model Open access Oct 2026

How well do large language models answer postoperative lumbar fusion questions? A blinded comparative analysis of accuracy, quality, readability, and safety-related content

Patients recovering from lumbar fusion increasingly seek guidance from large language model (LLM) chatbots when their surgeon is unavailable, but the accuracy, completeness, and safety-related content of such responses have not been compared across models for the postoperative period. To compare the accuracy...

Ahmet Kürşat Kara, Ozan Işık · 0 citations
Open access Sep 2026

Performance of large language models in postoperative hip fracture rehabilitation counseling for older adults: a comparative evaluation of safety, accuracy, reliability, readability, and empathy

To compare the performance of five widely used large language models in answering public questions about postoperative rehabilitation after hip fracture in older adults, focusing on safety, accuracy, reliability, readability, empathy, and overall information quality. This cross-sectional comparative evaluati...

Gen-Rong Hu, Qiu-Ju Lai, Shu-Qiong Liu et al. · 0 citations
Review Open access Aug 2026

Evaluation and Comparison of Large Language Model Responses to Frequently Asked Questions Regarding Patellofemoral Pain Syndrome: A Quality and Readability Assessment Study

The evaluated LLMs showed small and method-dependent differences in DISCERN-based information quality and variable differences in readability, but the findings should not be interpreted as evidence of factual accuracy, clinical safety, or suitability for individualized decision-making.

Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.