Skip to content
Open access

Evaluation and comparison of large language model responses to patient questions after diagnosis of high-risk human papillomavirus infection: an expert-rated digital patient education study

Jul 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 28 references
Medicine

TL;DR

ChatGPT-5.5 Instant achieved higher ratings in several quality domains, but response-level binary safety differences were imprecise and not statistically significant, but both models require guideline-based clinical oversight.

Abstract

Background After diagnosis of high-risk human papillomavirus (HPV) infection, patients often seek guidance on cancer risk, colposcopy, follow-up, treatment, partner management, pregnancy, and anxiety. Large language models (LLMs) are increasingly used for medical consultation, but their response quality and safety in this setting require evaluation. Methods In this single-center, expert-rated methodological study, 15 patient-style questions based on common outpatient consultations were submitted to ChatGPT-5.5 Instant and DeepSeek-V3. Each model generated three independent responses per question, yielding 90 artificial intelligence (AI)-generated responses. Five blinded gynecology experts independently evaluated all responses, producing 450 expert-rating records. Accuracy, safety, guideline concordance, completeness, and understandability were rated on a 1–5 Likert scale. Experts also assessed major error and potential harm. The primary outcome was the proportion of responses without major error or potential harm. Results ChatGPT-5.5 Instant had higher response-level scores than DeepSeek-V3 for accuracy (mean difference, 0.25; 95% CI, 0.11–0.39; FDR-adjusted p = 0.010), safety (0.27; 0.07–0.47; p = 0.010), completeness (0.22; 0.07–0.37; p = 0.020), and composite score (0.18; 0.04–0.31; p = 0.027). The difference in guideline concordance did not remain significant after multiplicity correction (0.16; −0.01 to 0.33; FDR-adjusted p = 0.057), and understandability was similar. Responses without major error or potential harm occurred in 42/45 (93.3%; 95% CI, 81.7–98.6%) ChatGPT responses and 38/45 (84.4%; 70.5–93.5%) DeepSeek-V3 responses (risk difference, 8.9 percentage points; 95% CI, −13.3 to 31.1; p = 0.344). Exploratory question-level risks were observed in HPV16/18 positivity, normal cytology, partner management, and pregnancy scenarios. Conclusion Both two models demonstrated generally strong performance across the evaluated quality domains. ChatGPT achieved higher ratings in several quality domains, but response-level binary safety differences were imprecise and not statistically significant. Both models require guideline-based clinical oversight.

Read PDF

Similar papers

Open access Sep 2026

Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study

Current mainstream LLMs demonstrate initial potential for generating educational content on bladder cancer, albeit with considerable heterogeneity across models, and advocate for a prudent, assistive role of LLMs in health communication under a human-AI collaborative model.

Dian Wan, You-Wen Li, Zheng Dong et al. · 0 citations
Sep 2026

Evaluation of the accuracy and reproducibility of large language models (ChatGPT, DeepSeek, Gemini) in responding to patient-centered lipedema questions.

LLMs can provide generally accurate and consistent responses to patient-centered questions about lipedema, particularly in areas related to general information and diagnosis, however, reduced accuracy and reproducibility in complex clinical domains suggest that expert oversight is essential when using these tools for p...

Rabia Sanır, E. Türkmen, E. Giray et al. · 0 citations
Open access Sep 2026

Safety-Oriented Benchmarking of Large Language Models in Risk-Based Management of Abnormal Cervical Screening Results: Scenario-Based Benchmark Study.

BACKGROUND Large language models (LLMs) are increasingly being considered for clinical decision support, yet their safety in risk-based cervical screening management remains insufficiently characterized. OBJECTIVE This study benchmarked the guideline concordance and safety-related performance of 3 LLMs in the initial...

Ö. Eroğlu, Cansın Eroğlu · 0 citations
Open access Sep 2026

Laboratory Medicine Decision Support—Beyond Exam Passing: A Blinded 100-Case Text-Based Benchmark of Diagnostic Accuracy, Management Quality, and Safety for ChatGPT, Gemini, and DeepSeek—LLM Decision Support in Laboratory Medicine

ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy.

K. Ulutaş, A. Pekmezci · 0 citations
Review Open access Aug 2026

Benchmarking publicly accessible large language models for English-language patient-facing acute pancreatitis information: a cross-sectional study of quality, transparency, and readability

Background Patients increasingly rely on large language models (LLMs) for health information, yet their suitability for decision-critical conditions such as acute pancreatitis remains unclear. Given that acute pancreatitis requires timely symptom recognition, severity assessment, treatment decision-making, recurrence p...

Biao Jiang, Hong-Xin Sun, Linlin Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.