Evaluating the accuracy of artificial intelligence-generated patient counseling for vestibular schwannoma.
Abstract
Objective
To compare the quality of Google, ChatGPT, and providers in answering common questions regarding vestibular schwannoma.
Methods
Search engine frequency data was used to generate ten common patient questions regarding the etiology, diagnosis, and management of vestibular schwannoma. Answers to these questions were generated from compiled Google search, ChatGPT-4, and provider responses. Response quality was graded by three blinded, independent, board-certified otolaryngologists using the DISCERN instrument. Readability scores were calculated. Inter-rater agreement was determined using Fleiss's kappa.
Results
Inter-rater agreement was high with a Fleiss's kappa score of 0.73. Total DISCERN scores among Google, ChatGPT, and expert response were 23.7, 30.4, and 28.4, respectively, with a standard error of 0.85. Average word count for ChatGPT (214 ± 98.2) responses were significantly longer (p < 0.05) than both Google (68.8 ± 72.3) and provider responses (108 ± 91.7) responses. ChatGPT had the highest readability score with a Flesch Reading Ease score of 33.1 ± 10.9, followed by Google (30.3 ± 20.7) and provider response (15.8 ± 8.70), with ChatGPT and Google performing significantly better than providers (p < 0.05).
Conclusion
The overall quality of ChatGPT responses to common questions about vestibular schwannoma was comparable to that of providers. Moreover, these responses were calculated to represent the highest readability levels when compared to both providers and Google search derived information. LEVEL OF EVIDENCE III.