Aug 2026· Otolaryngology Head & Neck Surgery· 0 citations· 25 references
Medicine
TL;DR
AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education, but they continue to lack guaranteed, verifiable sourcing.
Abstract
Objective
To evaluate the accuracy, completeness, clarity, source transparency, and readability of leading AI chatbot responses to patient questions about tracheostomy and to determine whether AI tools can reliably support patient education where high-quality guidance is critical for safety.
STUDY
Design
Cross-sectional content analysis.
Setting
Virtual study environment using publicly accessible AI platforms, with expert evaluation conducted via Qualtrics-based distribution.
Methods
Twelve frequently asked questions about tracheostomy care were identified using search-listening tools and clinician input, then submitted to 5 AI chatbots - ChatGPT4, Google Gemini 2.0, Microsoft Copilot, DeepSeek V3, and Grok 3 - and to a senior laryngologist. Three blinded laryngologists independently evaluated each response using the Quality Analysis of Medical Artificial Intelligence instrument. Readability was assessed using nine metrics.
Results
Gemini 2.0 achieved significantly higher completeness scores than physician responses (P < .001), with DeepSeek and Grok 3 (P < .05) also outperforming (P < .05). Accuracy did not differ significantly between AI- and expert-generated responses. On average, the AI models outperformed physician in clarity, completeness, and usefulness based on QAMAI scoring (P < .05). All AI and expert responses exceeded the NIH-recommended 6th-grade reading level, ranging from 10th-13th grade (P < .001). Inter-rater reliability was 78%.
Conclusion
AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education. However, they continue to lack guaranteed, verifiable sourcing, and this study did not assess actual patient comprehension of the AI-generated responses. Future efforts should focus on adapting AI-generated education materials to meet health literacy standards and evaluating their direct impact on patient understanding and outcomes.
Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...
A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al.· BMC Medical Informatics and...· 0 citations
ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery.
Kamil Balaban, Mehmet Batu Ertan, Mahmut Kalem· Digital Health· 0 citations
Patient comprehension of surgical information is often limited by medical jargon, health literacy barriers, and time constraints. Large language models (LLMs) such as ChatGPT‐5, Claude‐4, and Google AI search offer interactive context specific dialogue that may aid to overcome these limitations. To date, no study has...
Darcy Noll, T. Milton, P. Stapleton et al.· Trends in Urology & Men'...· 0 citations
DeepSeek demonstrates a significant advantage in information reliability, particularly excelling in postoperative and follow-up management content, and ChatGPT shows a slight edge in the readability of surgical planning sections.
Ling Tian, Long-Yue Tang, Ming-Tao Yang et al.· Frontiers in Public Health· 0 citations
AI-generated PILs offer brevity but do not consistently improve readability, with some indices suggesting increased complexity, and clinician oversight and further validation are essential to ensure AI-generated materials enhance, rather than hinder, patient understanding and engagement.
N. Kharma, S. Gill, Chayan Shanmugaratnam et al.· Journal of Orthopaedic Surge...· 0 citations
Background: Although artificial intelligence (AI) has been used in patient education for some time, the accuracy, reliability, and clinical appropriateness of AI-generated medical content remain inadequately defined and continue to be debated. This study aimed to evaluate the reliability, readability, and comprehensibi...
Furkan Türkoğlu, Elif Nur Gencer, Emre Erdoğan· Archives of Current Medical...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.