Skip to content

Performance of Large Language Model-Based Chatbots in Primary Health Care Teleconsultations: A Comparison Between Human and Artificial Intelligence-Generated Responses.

Jul 2026 · Telemedicine journal and e-health · pp. 15305627261471706 · 0 citations · 35 references
Medicine

TL;DR

AI demonstrates substantial potential as a support tool for teleconsultation services despite existing limitations, according to quality criteria, conciseness, coherence, and comprehensibility.

Abstract

INTRODUCTION Telehealth is a strategic component of primary health care and has advanced in Brazil through the National Telehealth Program. Its benefits can be enhanced by artificial intelligence (AI), which has emerged as a promising tool. This study aims to compare the performance of real human and AI-generated responses to queries submitted to the teleconsultation services of the Telehealth Center of the UFMG Faculty of Medicine (NUTEL FM-UFMG), a member of the Telehealth Brazil Program.

Methods

This is a comparative cross-sectional study of 180 real human and AI-generated responses, evaluated in a blinded manner according to quality criteria (medical adequacy, conciseness, coherence, and comprehensibility), risk potential, authorship identification accuracy, and inquiry resolution. Data from NUTEL FM-UFMG (January 2020 to May 2024) were utilized, covering cardiology, endocrinology, and obstetrics/gynecology (OB-GYN). Statistical analysis included the Shapiro-Wilk test, Kruskal-Wallis test, Nemenyi multiple comparison test, chi-square test, and Fisher's exact test.

Results

Across all specialties, a significant difference was observed in comprehensibility, with AI mean scores surpassing those of humans. For the remaining quality criteria, as well as for risk potential and inquiry resolution, no significant differences were found, despite AI scoring higher than humans. Within specific specialties, significant differences was observed in endocrinology (except conciseness) and cardiology (in conciseness); AI showed superior means. Across all specialties, as well as individually within endocrinology and OB-GYN, the accuracy of authorship identification (human vs. AI) was statistically significant.

Conclusion

Despite existing limitations, AI demonstrates substantial potential as a support tool for teleconsultation services.

View source

Similar papers

Review Open access Aug 2026

Multilingual Conversational AI Chatbots for Efficient Healthcare Delivery During Case History-Taking: A Systematic Review

Multilingual AI chatbots demonstrate a boost in healthcare efficiency, a reduction in language barriers, and the promotion of health equity, but exhibit challenges regarding validation, workflow integration, and evaluation standards, along with ethical issues such as privacy and bias.

R. Sharanesha, Deepti Virupakshappa, A. Abushanan et al. · 0 citations
Review Open access Aug 2026

A real‐world analysis of AI chatbot performance for medicines information enquiries

The study results demonstrated the chatbots' responses were typically suboptimal; albeit, a significant minority prompted a follow up to review the chatbot response, and the study conforms with the Declaration of Helsinki.

Duncan Yorkston, Tracey Borrie, Paul K. L. Chin · 0 citations
Open access Jul 2026

Evaluating Healthcare Provider and Artificial Intelligence Chatbot Responses to Patient Messages from a Health System Using the CREATE TRUST Framework.

Integration of artificial intelligence chatbots into healthcare requires rigorous, patient-centered evaluation. This study implements the CREATE TRUST framework-a novel tool evaluating both clinical substance and communication style-to compare responses from healthcare providers (HCPs) and two AI chatbots (GPT-4, Mixtr...

Matthew R. Allen, V. Tiyyala, Armaan S Johal et al. · 0 citations
Review Open access Aug 2026

Evaluation of generative AI-driven chatbots as sources of consumer health information on hand, foot, and mouth disease: a cross-sectional comparative study of safety, accuracy, information quality, readability, and empathy

Across a limited set of standardized English-language prompts, the five chatbots often generated coherent HFMD information, but occasional safety-relevant inaccuracies and readability barriers remained, suggesting these systems may support general information seeking but should not replace individualized professional a...

Zhao-Le Gong, Yan Na, Yi Guo et al. · 0 citations
Open access Oct 2025

Evaluating Frontline Health Workers' Responses to Patient Inquiries With and Without Large Language Model Support in Nigeria: Observational Study.

BACKGROUND Frontline health workers (FLWs) in low- and middle-income countries (LMICs) often face barriers that compromise care quality, including limited training, high patient loads, inadequate access to updated clinical guidance, and resource constraints. Working in underserved settings further exacerbates challenge...

Maria Moosa, Oluwaseyi Malumi, Solomon Chinedu et al. · 0 citations
Review Open access Feb 2026

Methods of Evaluating Large Language Model–Based Health Care Applications Used by Nonprofessionals: Protocol for a Scoping Review

A scoping review that maps approaches for evaluation of LLM-based applications used for health purposes by nonprofessionals and will provide directional guidance for further research and development in the field of quality assurance for LLM-based applications used by nonprofessionals is outlined.

Maren Keuchel, Pinar Bisgin, Tom Strube et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.