Skip to content

Large Language Models in Oral and Maxillofacial Surgery Triage: A Scoping Review

Aug 2026 · Current Surgery Reports · Vol 14 · 0 citations · 11 references

TL;DR

Large Language Models show potential in their diagnostic accuracy and consequent ability to reduce clinician burden, and may provide the greatest benefit when used to optimise referral quality at source, improving both clinician and potentially LLM triage downstream.

View source

Similar papers

Open access Jul 2026

A comparative evaluation of large language models in diagnosis and treatment planning in restorative dentistry.

OBJECTIVE With the advancement of artificial intelligence (AI), large language models (LLMs) have become an alternative source of information in dentistry. These LLMs, which can be trained on large data sets, can answer medical questions and provide references, but they can pose problems in terms of ethics and accurate information. This research aims to evaluate the accuracy of five different LLMs in diagnosis and treatment planning in the field of restorative dentistry. METHODOLOGY The 20 most common cases encountered in a restorative dentistry clinic were formulated into questions. The validity of the questions was assessed using the Lawshe Content Validity Index. The questions were posed to five different LLMs: ChatGPT-5, Deepseek V3.2, Claude Sonnet 4.5, Microsoft Copilot, and Google Gemini 3 Flash. Each model was asked to create a diagnosis and treatment plan for each case. The responses were evaluated by 42 restorative dentistry specialists using a Likert Scale. Additionally, the accuracy of the references provided in the responses was evaluated by the article authors. The obtained data were analyzed using the non-parametric Kruskal-Wallis test, and the Dunn multiple comparison test was applied in cases where significant differences were detected. RESULTS Statistically significant differences were found between the models for 15 out of 20 questions (p < 0.05). A significant difference was also found in terms of total median scores (p < 0.001), with Google Gemini 3 Flash (median:83) and ChatGPT-5 (median:81) achieving the highest scores. Reference quality was evaluated using a four-category framework. Claude Sonnet 4.5 and Google Gemini 3 Flash demonstrated the highest proportions of accurate and relevant references (85.7% and 84.2%, respectively), while DeepSeek V3.2 exhibited the highest fabrication rate (55.6%). CONCLUSIONS Based on specialist-evaluated response quality, no model demonstrated consistent and superior performance across all clinical scenarios. LLMs appear to have potential as supplementary information resources for clinicians in restorative dentistry; however, their clinical integration, impact on patient outcomes, and real-world usability remain to be established in future research.

Ebru İrem Teke, A. Borsöken · 0 citations
Review Open access Jul 2026

Patients' perception towards large language models in otorhinolaryngology, head and neck surgery: a single-centre survey

Objectives Large language models (LLMs) are increasingly discussed for use in clinical practice. Beyond their performance, patients' acceptance is crucial for their implementation. We investigated ORL-HNS patients' familiarity with AI/LLMs, use patterns, and trust in LLM-based medical information and recommendations. Methods In this single-centre prospective survey at a German university hospital, ORL-HNS patients with and without malignant disease completed a 15-item questionnaire. Results A total of 123 patients, 20 (16%) with and 103 (84%) without malignant disease, participated in the study. Most patients were familiar with the term AI (96%, n = 118) and LLMs (78%, n = 96). Overall, 72/123 (59%) reported using LLMs. One third (33%, n = 40) retrieved “Health information”, rating the LLMs with median Likert scores for comprehensibility 5 [IQR 4, 6], conciseness 5 [IQR 3, 6] and coherence 5 [IQR 3, 6]. However, perceived medical accuracy received a median rating of 4 [IQR 3, 5], significantly lower than comprehensibility (p < 0.05). With respect to the confidence in the recommendations exclusively by LLMs [median 2 (IQR 2, 3.5)] received significantly lower ratings than doctors [median 5 (IQR 5, 6)] and doctors also using LLMs [median 5 (IQR 4, 6)], p < 0.0001 respectively. Conclusion ORL-HNS patients are largely familiar with LLMs and frequently use them, but their trust and confidence regarding health information provided by LLMs alone is limited. Patients show the greatest confidence in doctors' recommendations. Yet they reported similar confidence in physician recommendations and physician recommendations supported by LLMs, suggesting that clinician-led LLM use may be acceptable to many patients.

C. Buhr, A. Blaikie, Harry Smith et al. · 0 citations
Open access Jul 2026

Language-dependent performance variation in large language models for dental trauma management: a comparative evaluation of ChatGPT-5.2, Gemini 3.0, and Claude 4.5 Sonnet.

BACKGROUND Large language models (LLMs) are increasingly evaluated for medical question answering and clinical information tasks, yet the impact of query language on their performance in specialized domains such as dental traumatology remains insufficiently studied. The primary objective was to evaluate whether query language (English vs. Turkish) affects LLM performance in a controlled scenario-based assessment of dental trauma management. Secondary objectives were to compare overall performance across three LLMs and to examine whether language effects are uniform across models or model-specific. METHODS Twenty-seven clinical scenarios covering 13 dental trauma categories were presented to ChatGPT 5.2, Gemini 3.0, and Claude 4.5 Sonnet in both English and Turkish, generating 162 responses. Two blinded endodontists independently evaluated responses using a standardized rubric assessing accuracy (40%), completeness (35%), and safety (25%) against IADT 2020 Guidelines. Inter-rater reliability was assessed using intraclass correlation coefficient (ICC). Language effects were analyzed using Wilcoxon signed-rank tests; model comparisons employed Kruskal-Wallis and Mann-Whitney U tests with Bonferroni correction. RESULTS Inter-rater reliability ranged from moderate to good across evaluation dimensions (ICC: 0.738-0.836). ChatGPT showed the strongest language effect with 9.14% higher performance in English (p < 0.001, r = 0.874). Gemini showed moderate English advantage (5.69%, p = 0.003, r = 0.572). Claude exhibited language independence with virtually identical performance in both languages (-0.02%, p = 0.220). In English, significant model differences emerged (H = 22.31, p < 0.001); however, model performance converged in Turkish (H = 2.89, p = 0.236). CONCLUSIONS This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management. ChatGPT 5.2 achieved the highest performance in English but exhibited the most pronounced Turkish-language degradation, including substantial safety score decline. Gemini 3.0 showed an intermediate pattern with moderate English advantage. Claude 4.5 Sonnet demonstrated language-independent performance across all evaluated dimensions. These findings are based on a standardized scenario-based assessment and should not be extrapolated to real clinical environments or patient care settings.

H. Öz, M. Dundar · 0 citations
Jul 2026

Diagnostic accuracy of large language models in ICOP-based orofacial pain diagnosis: A comparative study.

OBJECTIVE To compare the diagnostic performance of ChatGPT 5.5, Claude Opus 4.1, Gemini 3 Flash, and Grok 4 in International Classification of Orofacial Pain (ICOP)-based clinical scenarios. METHODS Thirty ICOP diagnoses were randomly selected, and corresponding clinical scenarios were manually developed. Each scenario was submitted to all models using standardized prompts in independent sessions. Two blinded evaluators assessed primary diagnosis accuracy, subclassification accuracy, clinical interpretation, and management recommendations. RESULTS  Overall performance differed significantly among models (p < .001). Grok 4 achieved the highest total score and outperformed the other models. No significant differences were found among ChatGPT 5.5, Gemini 3 Flash, and Claude Opus 4.1. Subclassification accuracy was consistently lower than primary diagnosis accuracy, while management recommendations did not differ significantly. CONCLUSION LLM performance varied across ICOP-based scenarios. Although Grok 4 showed the highest diagnostic concordance, current LLMs should support, not replace, clinician judgment.

M. S. Şimşek, Enis Esen, M. Koparal · 0 citations
Review Open access Aug 2026

Evaluation of a large language model for clinician-facing preoperative cost-communication preparation in total knee arthroplasty

Background and aims Costs associated with total knee arthroplasty (TKA) may affect treatment preparation, expectation management, and postoperative care planning. Previous large language model (LLM) studies have focused mainly on medical question answering, patient education, and clinical decision support, whereas their performance in clinician-facing preoperative cost-communication preparation remains unclear. This study evaluated an LLM using real-world clinical records from two hospitals in an expert-referenced offline evaluation. Methods Preoperative medical records of 80 patients who underwent primary unilateral TKA at two hospitals in China from January to May 2026 were included. A structured expert-panel process was used to develop a preoperative cost-communication framework comprising 4 dimensions and 17 clinical cues and to establish case-level minimum necessary communication items (CL-MNCIs) for each case. Task 1 assessed identification of the 17 cues. Task 2 assessed CL-MNCI coverage and classified all generated items according to case relevance, medical-record support, redundancy, and safety. Twenty-four cases were non-randomly selected by CL-MNCI count for three repeated-generation runs. Results Task 1 comprised 1,360 case–label classification units. Micro-precision, micro-recall, micro-F1, accuracy, MCC, and macro-F1 were 0.901, 0.888, 0.894, 0.948, 0.860, and 0.840, respectively. Experts established 569 CL-MNCIs, of which 483 were covered, yielding an overall coverage rate of 84.9%; complete coverage was achieved in 17 cases. The LLM generated 716 items, including 12 safety events across 9 cases, 483 items matching CL-MNCIs, 132 record-supported supplementary items, 50 redundant items, and 39 items with insufficient record support. Overall, 665 items (92.9%) were case-relevant and record-supported, although this proportion included redundant content. Pairwise Jaccard similarity for covered CL-MNCI sets ranged from 0.813 to 0.823. Of 180 CL-MNCIs, 126 (70.0%) were covered in all three runs, and case-level agreement in safety classification ranged from 87.5 to 95.8%. Conclusion The LLM showed offline potential for identifying cost-communication cues and generating clinician-facing preparation checklists for TKA, but content omissions, quality variation, safety risks, and substantive cross-run variability remained. Its use should be limited to clinician-reviewed communication preparation and should not replace direct patient cost disclosure or professional judgment.

Zebing Ma, Liping Xue, Gonghui Jian et al. · 0 citations
Open access Mar 2026

Can Large Language Models Identify When an Upper Extremity Problem Needs Nonurgent Attention? An Assessment of Multiple LLM Chatbots.

IntroductionSeeking emergency care regarding musculoskeletal sensations is far more prevalent than limb or life-threatening pathophysiology. We studied the ability of an LLM to distinguish between urgent and nonurgent upper extremity symptoms and provide an accurate diagnosis.MethodsFive LLMs (ChatGPT-4, ChatGPT-4o, Co-Pilot, Gemini, and PerplexityChat) were presented with descriptions of seven urgent and seven nonurgent symptom scenarios written below a sixth grade reading level. LLM responses were identified as appropriate if immediate urgent medical attention was recommended after an initial and ongoing inquiry ("What additional information do you need to diagnose my condition?"). Diagnoses provided were identified as correct, partially correct, or incorrect. The analysis was repeated 24 months later with the current LLM versions and results were compared.ResultsLLMs discerned nonurgent conditions with 97% positive predictive value (PPV) and an 89% negative predictive value (NPV) on initial query, which improved to 96% and 97% respectively after ongoing inquiry. Compartment syndrome was misidentified as nonurgent in 80% of scenarios on initial inquiry, although four of five LLMs corrected on continued inquiry. Diagnosis was correct or partially correct for 115 of 150 (82%) on initial inquiry. An updated analysis 24 months later demonstrated marked improvement in LLM ability to identify emergencies with 100% PPV and 95% NPV on initial query and 98% NPV after ongoing inquiry.ConclusionThe finding that LLMs can distinguish urgent from nonurgent upper extremity conditions suggests that artificial intelligence tools could help reduce unnecessary use of high-cost emergency services, allowing those resources to be reserved for patients who require timely care.

Jefferson Hunter, David Ring, Prakash Jayakumar · 0 citations

Related blog posts

Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.