May 2026· arXiv.org· Vol abs/2605.28025· 1 citation· 48 references
Computer Science
TL;DR
The Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals, is introduced.
Abstract
Existing safety evaluations for large language models overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answered all medical questions, but responses to low health-literacy signals consistently omitted more key information, provided fewer concrete next steps, and offered less support for independent judgment. We term this pattern Differential Information Dilution (DID). A comparison with 300 real-world health queries provides preliminary evidence of rank-order validity. A knowledge-guided mitigation prompt reduces information dilution for most models, with the largest reductions in underinformative simplification observed for Claude (~8%) and Qwen (~6%). Code and data are available at https://github.com/Rainxu09/MIRA.
Large Language Models (LLMs) are increasingly used by consumers as sources of health information. Evaluating the quality of such systems requires considering multiple quality dimensions rather than relying on a single indicator. Response consistency is often interpreted as a sign of reliability, but it remains unclear...
Line Praestegaard, Elda Paja· 2026 IEEE 34th International...· 0 citations
It is concluded that neither consistency nor adaptation can currently be considered clearly preferable, highlighting the need for empirical evidence on which approach better serves users across cultural contexts.
M. Bui, Mario Sanz-Guerrero, Abteen Ebrahimi et al.· 0 citations
A side effect that misrepresents patients is measured: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes the question never stated, in effect rewriting who the patient is.
BACKGROUND
large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality,...
Jun-Zheng Li, Ying-Jie Wu, Man Yang et al.· Nutrición Hospitalaria· 0 citations
Public-facing chatbots may support general patient education but should not replace individualized clinician-led prognostic communication as public-facing chatbots differed substantially in safety, reliability, communication quality, and readability.
Yan-Ru Jiang, Qian-Yun Wang, Liang Zheng et al.· Frontiers in Public Health· 0 citations
Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated,...
J. Cahoon, C. Stanwyck, Sulaiman Somani et al.· 0 citations