Skip to content
Open access

Accuracy and response repeatability of three large language models on undergraduate operative dentistry multiple-choice questions

Aug 2026 · Frontiers in Dental Medicine · Vol 7 · 0 citations · 17 references
Medicine

TL;DR

All three large language models demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool, however, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.

Abstract

Introduction Artificial intelligence (AI) has emerged as a valuable tool in the field of dentistry to assist healthcare providers in diagnosis, optimize treatment outcomes, support research, and improve education. Furthermore, the generative aspects of AI have emerged as a powerful tool for dental professionals to tackle clinical and academic challenges. However, the reliability of AI in domains such as undergraduate training remains unverified. Aim This study aimed to evaluate the accuracy and consistency of three large language models (LLMs) in answering multiple-choice questions on the undergraduate operative dentistry curriculum to determine their reliability as a supplementary learning resource. Methods Sixty multiple-choice questions were formulated from the undergraduate operative dentistry curriculum. Each LLM (Claude, Gemini, and ChatGPT) was queried individually nine times over three days. The results were evaluated using a predetermined answer key. The accuracy of the LLMs was compared using a generalized estimating equation (GEE) with question-level clustering analysis. Consistency was assessed at the question level using response consistency, correctness consistency, and Fleiss' kappa. Statistical significance was set at P < 0.05. Results Gemini demonstrated the highest overall accuracy (98.89%), followed by ChatGPT (98.33%) and Claude (97.78%). No statistically significant difference in accuracy was observed among the three LLMs (GEE, overall Wald χ2(2) = 1.46, p = 0.481). All three models showed almost perfect intersession agreement (Fleiss' κ = 0.93–0.95), with response and correctness consistency ranging from 88.3% to 91.7% across models. Conclusions All three LLMs demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool. However, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.

Read PDF

Similar papers

Open access Sep 2026

Evaluating the Accuracy of Large Language Models in Dentistry: A Multi-Model Study Using Clinical Questions From Turkey's Dental Specialty Exams.

While LLMs show strong potential in supporting dental education through standardised exams, their performance varies by model and question type, and further improvements are needed to enhance reliability across different dental disciplines.

N. Acar, Fatih Sengul, Periş Çelikel et al. · 0 citations
Open access Aug 2026

Accuracy and Consistency of Artificial Intelligence Chatbots in Dental Anatomy Education: A Comparative Study

The potential of Artificial Intelligence (AI), large language models (LLMs) in enhancing dental education emphasises the need for careful selection of AI tools to improve learning outcomes. Therefore, this study evaluates the accuracy and consistency of responses from eight AI chatbots to multiple‐choice questions...

M. B. Mirza, A. Robaian, A. Alqahtani et al. · 0 citations
Sep 2026

Artificial Intelligence in Evaluating Undergraduate Dental Students: Benchmarking AI Platforms Against Student Performance.

INTRODUCTION Artificial intelligence (AI), particularly large language models (LLMs), has demonstrated significant potential in educational assessment across medical disciplines. This study explores whether such models can match or surpass the performance of undergraduate dental students in a rigorous academic examinat...

Alexia-Ecaterina Cârstea, V. Vasilescu, Ana-Maria Cristiana Țâncu et al. · 0 citations
Sep 2026

Comparative evaluation of large language models in interpreting the scientific literature on intraoral scanners across varying input levels.

Evaluating the response accuracy of 2 advanced large language models found the Gemini-3.0 Flash responses excelled in highly structured, direct-response questions to relevance, clarity, and closeness to the article, whereas the ChatGPT-5.2 response demonstrated greater scientific accuracy at all levels and performed be...

Pranay Jain, F. R, A. V et al. · 0 citations
Review Open access Sep 2026

Assessing the accuracy and usability of artificial intelligence-based language models in responding to common periodontal patient questions.

BACKGROUND Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aim...

Sajad Jahantigh, R. Amid, A. Moscowchi et al. · 0 citations
Open access Aug 2026

Accuracy and Consistency of Three Large Language Models on Fixed Prosthodontics Questions

Large language models (LLMs) are increasingly used in dental education and clinical settings, but their accuracy and consistency in fixed prosthodontics remain uncertain. Objectives: To compare the accuracy (using a strict three-attempt criterion) and repeated response consistency of ChatGPT-5.4, Gemini 3.1 Pro, and Cl...

A. Bashir, Moeen Ud Din Ahmad, Ussamah Waheed Jatala et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.