Skip to content

MIRA: A Bilingual Benchmark for Medical Information Response Audit

May 2026 · arXiv.org · Vol abs/2605.28025 · 1 citation · 48 references
Computer Science

TL;DR

The Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals, is introduced.

Abstract

Existing safety evaluations for large language models overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answered all medical questions, but responses to low health-literacy signals consistently omitted more key information, provided fewer concrete next steps, and offered less support for independent judgment. We term this pattern Differential Information Dilution (DID). A comparison with 300 real-world health queries provides preliminary evidence of rank-order validity. A knowledge-guided mitigation prompt reduces information dilution for most models, with the largest reductions in underinformative simplification observed for Claude (~8%) and Qwen (~6%). Code and data are available at https://github.com/Rainxu09/MIRA.

View source

Similar papers

Conference Aug 2026

When Stable Answers are Not Enough: Evaluating Response Consistency and Medical Alignment in Consumer-Facing LLMs

Large Language Models (LLMs) are increasingly used by consumers as sources of health information. Evaluating the quality of such systems requires considering multiple quality dimensions rather than relying on a single indicator. Response consistency is often interpreted as a sign of reliability, but it remains unclear...

Line Praestegaard, Elda Paja · 0 citations
Open access Sep 2026

Performance of large language models in answering public questions about nutrition in cirrhosis: a comparative study.

BACKGROUND large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality,...

Jun-Zheng Li, Ying-Jie Wu, Man Yang et al. · 0 citations
Open access Aug 2026

Safety and quality of public chatbots for lung cancer prognostic information: a comparative evaluation

Public-facing chatbots may support general patient education but should not replace individualized clinician-led prognostic communication as public-facing chatbots differed substantially in safety, reliability, communication quality, and readability.

Yan-Ru Jiang, Qian-Yun Wang, Liang Zheng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Living Benchmark for Information Retrieval from Electronic Health Records

Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated,...

J. Cahoon, C. Stanwyck, Sulaiman Somani et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.