Skip to content
Open access

Quality, readability, and clinical-risk signals of default public-interface LLM responses to vascular and perioperative patient questions: a cross-sectional snapshot

Aug 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 32 references
Medicine

TL;DR

In this English-language benchmark of default first responses from five public LLM interfaces accessed through specific logged-in accounts from a Hong Kong, China IP address during a single late-May 2026 window, 107 of 110 responses received the maximum Guideline Concordance Score, and no response met the prespecified criteria for a high Potential Clinical-Risk Severity Flag.

Abstract

Background Patients increasingly use public large language model chatbot interfaces to seek health information. In vascular disease and perioperative management, default first responses may influence how patients interpret urgent symptoms, antithrombotic medications, procedural choices, and anesthesia-related safety issues. Methods This cross-sectional benchmark study evaluated 110 default first responses returned by five publicly accessible LLM chatbot interfaces during a defined access window on May 28–29, 2026, Beijing time (UTC + 8). Interface names were recorded solely as the public-interface display labels visible at the time of access and should not be interpreted as independently verified API-level model identifiers. Each model was queried with 22 guideline-derived English patient-facing questions, yielding 110 responses. Responses were assessed using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), Journal of the American Medical Association (JAMA) benchmark criteria, used only as a visible metadata/transparency proxy, six readability formulas, an investigator-developed Guideline Concordance Score, and an investigator-developed Potential Clinical-Risk Severity Flag. Differences across the evaluated public-interface response sets were tested using Friedman tests with Holm-adjusted post hoc comparisons. Results Interrater agreement was high for established instruments: DISCERN ICC(A,1) = 0.940, EQIP ICC(A,1) = 0.830, GQS weighted κ = 0.829, and JAMA weighted κ = 0.898. Observed public-interface response performance differed significantly for DISCERN, EQIP, and JAMA criteria (all p < 0.001), and for GQS (p = 0.011). Within this defined late-May 2026 response set, responses returned by the interface displaying the label ‘Grok-4.3’ had the highest observed sample means for DISCERN (58.68), EQIP (81.14), GQS (4.27), and the JAMA visible metadata/transparency proxy score (1.00). No evaluated public-interface response set achieved recommended sixth-grade readability, and no individual response met all six readability thresholds. Responses returned by the interface displaying the label ‘DeepSeek-v4’ had the most favorable observed readability profile within this sampled response set, although the mean FKGL remained 11.53. These observations should not be interpreted as durable model rankings. Conclusion In this English-language benchmark of default first responses from five public LLM interfaces accessed through specific logged-in accounts from a Hong Kong, China IP address during a single late-May 2026 window, 107 of 110 responses received the maximum Guideline Concordance Score, and no response met the prespecified criteria for a high Potential Clinical-Risk Severity Flag. Pronounced ceiling and floor effects preclude conclusions about clinical sufficiency or safety. These findings should not be generalized to non-English use, different health-literacy levels, country-specific emergency-care pathways, other regions or account configurations, or later interface states.

Read PDF

Similar papers

Review Open access Aug 2026

Benchmarking publicly accessible large language models for English-language patient-facing acute pancreatitis information: a cross-sectional study of quality, transparency, and readability

Background Patients increasingly rely on large language models (LLMs) for health information, yet their suitability for decision-critical conditions such as acute pancreatitis remains unclear. Given that acute pancreatitis requires timely symptom recognition, severity assessment, treatment decision-making, recurrence p...

Biao Jiang, Hong-Xin Sun, Linlin Chen · 0 citations
Open access Aug 2026

Safety, accuracy, empathic communication, information quality, and readability of five large language model interfaces answering public questions about interstitial cystitis/bladder pain syndrome

Publicly accessible LLM interfaces showed domain-specific differences when answering IC/BPS-related public questions and may support general IC/BPS education and question preparation but should not replace clinician-led evaluation or individualized medical advice.

Jiang-Tao Zhu, Zhen-Hua Zhao, Song Li et al. · 0 citations
Open access Sep 2026

Evaluating large language models for myocardial infarction public health education: a comparative study on information quality, transparency and readability

While LLMs can generate structurally clear and logically coherent foundational content for MI-related queries, they occasionally produce clinically inappropriate directives, creating substantial reading barriers for the general public.

Tai-Long Lv, Wen-Kai Bao, Shu-Di Li et al. · 0 citations
#large language models Review Open access Sep 2026

A cross-sectional evaluation of large language model chatbot interfaces for patient-facing herpes zoster information: safety, information quality, and readability

In this standardised benchmark, LLM chatbot interfaces provided generally favourable accuracy and information-quality scores but showed clinically relevant safety limitations, sparse source attribution and disclosure, and readability challenges.

Da-Lian Liang, Cong Mai, Zheng-Kun Zhang et al. · 0 citations
Review Open access Aug 2026

Quality, readability, and patient safety of ChatGPT-generated responses to fall-related questions in older adults: a multidisciplinary evaluation

While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns and AI-generated health content should be reviewed and tailored to older adults' health literacy needs before clinical use.

Merve Arı, N. Ilçin, Hatice Yağcıoğlu et al. · 0 citations
Open access Aug 2026

Evaluating search-enabled large language model interfaces for mpox public health consultation: a guideline-based comparative study

The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability, which support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are eva...

Qi-Qi Zheng, Ru Chen, Ming-Ming Cai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.