Skip to content
Review Open access

Quality, readability, and patient safety of ChatGPT-generated responses to fall-related questions in older adults: a multidisciplinary evaluation

Aug 2026 · Revista da Associação Médica Brasileira · Vol 72 · 0 citations · 20 references
Medicine

TL;DR

While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns and AI-generated health content should be reviewed and tailored to older adults' health literacy needs before clinical use.

Abstract

SUMMARY

Objective

Older adults increasingly use artificial intelligence-based tools to obtain health information. Although artificial intelligence chatbots such as ChatGPT may enhance access, the quality, readability, and patient safety of fall-prevention information remain uncertain. This study aimed to evaluate the quality, readability, and patient safety implications of ChatGPT-generated responses to common questions about fall risk and home safety in older adults.

Methods

Ten frequently asked fall-related questions were submitted to ChatGPT (version 5.2). Responses were independently assessed by a multidisciplinary panel including physiotherapists, a geriatrician, a physical medicine and rehabilitation physician, an occupational therapist, and an orthopedic specialist. Quality was evaluated using the Mika classification. Readability was measured with the Flesch-Kincaid Grade Level. Interrater reliability was analyzed using a two-way random-effects intraclass correlation coefficient model with absolute agreement (intraclass correlation coefficient [2,k]).

Results

Three responses were rated as "excellent," while seven responses were rated as "satisfactory requiring minimal clarification." No response received a rating corresponding to "moderately satisfactory" or "unsatisfactory." The mean Flesch-Kincaid Grade Level was 8.4 (range 4.3–11.9). Five responses exceeded the readability levels commonly recommended for patient education materials. Interrater reliability demonstrated fair agreement (intraclass correlation coefficient [2,k]=0.72; 95%CI 0.64–0.80).

Conclusion

While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns. AI-generated health content should be reviewed and tailored to older adults’ health literacy needs before clinical use.

Read PDF

Similar papers

Open access Jul 2026

ChatGPT-5 as a leisure health advisor: multidimensional assessment of reliability, quality, usefulness and readability

Background The demonstrated protective effects of leisure activities on physical and mental health underscore the need for accessible guidance. Large Language Models (LLMs) like ChatGPT-5 offer a potential solution, yet their application in non-clinical leisure health advice requires rigorous evaluation. This study aim...

Alican Bayram · 1 citation
Open access Aug 2026

Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study

ABSTRACT Objective: Large language model chatbots are increasingly consulted for medical information. This study evaluated the accuracy and readability of chatbot responses to common patient questions on dry eye disease.Methods: This cross-sectional study analysed responses from four chatbots (ChatGPT 3.5, ChatGPT 4.0,...

B. Kesimal, S. Kocamış · 0 citations
#small language model Open access Sep 2026

Safety, accuracy, empathy, information quality, and readability of publicly accessible LLM-based chatbots for traumatic brain injury and concussion questions: a cross-sectional comparative study

Background Large language model (LLM)-based chatbots are increasingly used by the public to obtain health information, but their performance in answering questions related to traumatic brain injury (TBI) and concussion remains unclear. This study evaluated five publicly accessible LLM-based chatbots across safety, accu...

Xin Zuo, Huan Zuo, Min Zhang et al. · 0 citations
Review Open access Jul 2026

Evaluating the reliability, quality, and readability of AI-generated patient education on hallux valgus: a comparative study of large language models

Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...

A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.