Skip to content
Review Open access

Benchmarking publicly accessible large language models for English-language patient-facing acute pancreatitis information: a cross-sectional study of quality, transparency, and readability

Aug 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 33 references
Medicine

Abstract

Background Patients increasingly rely on large language models (LLMs) for health information, yet their suitability for decision-critical conditions such as acute pancreatitis remains unclear. Given that acute pancreatitis requires timely symptom recognition, severity assessment, treatment decision-making, recurrence prevention, and follow-up management, LLM-generated information should demonstrate reliability, transparency, and readability. Objectives To evaluate the informational quality, visible transparency-related features, and readability of English-language responses generated by five publicly accessible LLMs to standardized patient-facing questions on acute pancreatitis. Methods This cross-sectional benchmark study developed 24 English-language, single-intent questions on acute pancreatitis across six clinical domains using public search intents and guideline-derived decision-critical content. The analysis focused exclusively on English-language patient-facing responses. Each question was submitted once to GPT-5.4 Thinking, DeepSeek-V3.2, Gemini 3.1 Pro, Grok 4.3, and Qwen3.6-Max-Preview, generating 120 responses. Anonymized responses were independently evaluated by two blinded gastroenterologists using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), and Journal of the American Medical Association (JAMA) benchmark criteria. Readability was assessed using six established formulas. A structured response-level safety analysis was added to evaluate factual inaccuracies, clinical hallucinations, and clinical safety-risk severity. Between-model differences were analyzed with Friedman tests followed by post hoc paired Wilcoxon signed-rank tests with Holm correction. Results Significant between-model differences were observed across all quality, transparency-related, and readability outcomes. Grok 4.3 achieved the highest mean DISCERN, GQS, EQIP, and JAMA scores, reflecting the strongest informational quality profile according to the predefined quality instruments, although visible transparency cues remained limited across all models. In the added safety analysis, factual inaccuracies and clinical hallucinations were each identified in 16 of 120 responses (13.3%), whereas clinical safety-risk signals were identified in 7 responses (5.8%), all of which were adjudicated as score 1 (low risk) under the predefined clinician-rated rubric; no moderate- or high-risk clinical safety event was adjudicated. DeepSeek-V3.2 demonstrated the most favorable readability profile, with the highest Flesch Reading Ease score (44.42 ± 12.27), which nevertheless remained substantially below the recommended threshold of ≥ 80. None of the 120 responses satisfied all six predefined readability targets. All model-level readability distributions differed significantly from recommended thresholds in the direction of poorer readability. Conclusion Publicly accessible LLMs generated English-language responses on acute pancreatitis with variable informational quality, limited visible transparency cues, and consistently inadequate readability under standardized default public-interface conditions. Because the primary analysis used a single-generation cross-sectional design, model rankings should be interpreted as performance snapshots rather than definitive or temporally stable hierarchies. Higher quality scores did not necessarily establish factual accuracy, clinical safety, patient-education-level readability, or stronger response-level transparency. Current public-interface LLM outputs may serve as clinician-reviewed drafts for patient education but should not function as standalone patient resources. Future AI-based health information systems should strengthen clinical completeness, plain-language communication, visible evidence support, actionability, and clinician-supervised safeguards.

Read PDF

Similar papers

Open access Aug 2026

Quality, readability, and clinical-risk signals of default public-interface LLM responses to vascular and perioperative patient questions: a cross-sectional snapshot

In this English-language benchmark of default first responses from five public LLM interfaces accessed through specific logged-in accounts from a Hong Kong, China IP address during a single late-May 2026 window, 107 of 110 responses received the maximum Guideline Concordance Score, and no response met the prespecified...

Wei Zhong, Yuanyuan Zhang, Yu Huang et al. · 0 citations
Open access Sep 2026

Evaluating large language models for myocardial infarction public health education: a comparative study on information quality, transparency and readability

While LLMs can generate structurally clear and logically coherent foundational content for MI-related queries, they occasionally produce clinically inappropriate directives, creating substantial reading barriers for the general public.

Tai-Long Lv, Wen-Kai Bao, Shu-Di Li et al. · 0 citations
Review Open access Sep 2026

Evaluating large language models as clinical decision support tools in primary healthcare settings: Protocol for a multi-country comparative validation study on expert-adjudicated hypothetical vignettes (hypMOOVE-PHC)

The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania, and aims to validate a pool of LLMs through clinical review of expert-generated vignettes through fully crossed repeated-measures comparative eva...

P. Macharia, C. Kachimanga, M. Mahende et al. · 0 citations
Open access Sep 2026

Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study

Current mainstream LLMs demonstrate initial potential for generating educational content on bladder cancer, albeit with considerable heterogeneity across models, and advocate for a prudent, assistive role of LLMs in health communication under a human-AI collaborative model.

Dian Wan, You-Wen Li, Zheng Dong et al. · 0 citations
Open access Jul 2026

Evaluation and comparison of large language model responses to patient questions after diagnosis of high-risk human papillomavirus infection: an expert-rated digital patient education study

ChatGPT-5.5 Instant achieved higher ratings in several quality domains, but response-level binary safety differences were imprecise and not statistically significant, but both models require guideline-based clinical oversight.

Zhen Hao, Lin Wang, Yue Wu et al. · 0 citations
Review Open access Sep 2026

Large language models for patient-facing pathology report interpretation: A scoping review.

LLM-based patient-facing pathology report interpretation shows potential to bridge specialist pathology language and patient communication, and evaluation should extend beyond readability to include fidelity to the original pathology report, patient understanding, safety, and usability.

Chen Wang, Jie Hao, Si-Jia Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.