Skip to content

Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models

Jul 2026 · arXiv.org · Vol abs/2607.27384 · 0 citations · 16 references
Computer Science

TL;DR

NarrativeShield, a three-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near-zero and achieving the lowest rate of severely unstable decisions of any method across all models, at a modest and mechanistically expected accuracy cost.

Abstract

Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge. Unlike prior demographic-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of variation, with no demographic marker present in any form. We construct a dataset of 1,000 USMLE clinical vignettes, each rewritten into three sociolinguistically distinct personas under an independently audited fact-preservation guarantee, verified by a separate model that never sees the generation prompt. Across seven language models spanning three architecture families and scales, Narrative Anchoring is statistically significant under direct prompting in every model tested, with a Narrative Anchoring Gap of 0.064 to 0.151. Chain-of-thought reasoning and explicit debiasing instructions reduce the bias only partially, and their apparent gains are frequently confounded by accuracy collapse. We introduce NarrativeShield, a three-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near-zero ($-0.004$ to $0.037$) and achieving the lowest rate of severely unstable decisions (DSS $<$ 0.8) of any method across all models, at a modest and mechanistically expected accuracy cost for most models. A stress test using a non-instruction-tuned base model shows that executing a debiasing intervention at all is gated by zero-shot instruction-following ability, not prompt content alone. We release our dataset, human-validated for fact preservation, as a standalone resource for studying register-based clinical bias.

View source

Similar papers

Preprint Aug 2026

Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over.

Augusto Bernardo Pissarra, Victor Farias DE Souza · 0 citations
#artificial intelligence Preprint Aug 2026

Generating Clinical Vignettes that Preserve Cognitive Formulations

Results show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation and show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation.

Amit Oren, N. Hertz-Palmor, Dean Ariel et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models

Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs) are increasingly queried to interpret images in ways that touch on medical or diagnostic judgments, raising safety concerns when such infere...

Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel et al. · 0 citations
Preprint Aug 2026

CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives

This work proposes CRAFT, an LLM framework that pairs a generator with a constraint-based verifier to iteratively produce and refine stage-wise symptom timelines through targeted feedback, and demonstrates that CRAFT consistently improves temporal ordering accuracy.

Chengyang He, Tahreem Arif, Marko Zivkovic et al. · 0 citations
Open access Aug 2026

Poster 162. Racial Biases Perpetuated by Modern Large Language Models Negatively Impact Diagnostic Reasoning and Treatment Recommendations in Musculoskeletal Healthcare

Objectives: To determine whether contemporary large language models (LLMs) perpetuate gender and racial biases in medical decision-making. Methods: A total of 180 standardized vignettes concerning musculoskeletal diagnoses and treatments were extracted from AAOS Restudy and Orthobullets. A customized GPT-o4-mini model...

Kyle N. Kunze, Nicholas Allen, Sophia J. Madjarova et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.