Skip to content

Sociodemographic bias in LLMs' clinical decision-making for dizziness.

Aug 2026 · Journal of Vestibular Research-Equilibrium & Orientation · pp. 9574271261474218 · 0 citations · 16 references
Medicine

TL;DR

LLM clinical recommendations varied by sociodemographic descriptors, particularly under diagnostic uncertainty, suggesting structured inputs may mitigate bias in clinical AI systems.

Abstract

ObjectiveAs large language models (LLMs) enter clinical decision support, concerns persist about sociodemographic bias. We assessed whether LLM recommendations for dizziness vary by patient descriptors and clinical detail.MethodsWe conducted a cross-randomized in-silico vignette study. One hundred synthetic emergency department dizziness cases were created using established diagnostic frameworks including the TiTrATE paradigm, SAEM GRACE-3 guidelines, and Bárány Society diagnostic criteria. Each vignette was tested in a neutral form and with 33 sociodemographic descriptor variants (34 total). Twelve instruction-tuned LLMs from multiple model families were evaluated. Models answered five binary clinical decision questions addressing etiology classification, triage disposition, neuroimaging, bedside vestibular examination, and mental health referral. Each model-vignette-descriptor combination was repeated 10 times, yielding 2,040,000 responses. Sociodemographic bias was quantified as descriptor-specific percentage-point deviations from neutral control recommendations with 95% confidence intervals.ResultsSociodemographic descriptors influenced LLM recommendations, with the largest differences observed for mental health referral decisions in diagnostically ambiguous cases. Referral likelihood was lower for Black transgender women (-12.2 pp; 95% CI -14.0 to -10.3), Black patients experiencing homelessness (-9.1 pp; -11.0 to -7.3), and patients experiencing homelessness (-7.7 pp; -9.5 to -5.9). Differences were attenuated when vignettes contained clearer diagnostic information. Other effects were smaller, including increased neuroimaging recommendations for low-income descriptors (+4.0 pp; 95% CI 2.1-5.8).ConclusionLLM clinical recommendations varied by sociodemographic descriptors, particularly under diagnostic uncertainty. More detailed clinical information reduced these disparities, suggesting structured inputs may mitigate bias in clinical AI systems.

View source

Similar papers

Open access Sep 2026

Performance comparison of large language models in interpreting clinical guidelines for migraine prevention: A multidimensional analysis

While all four models exhibited comparable adherence to guideline-based content, DeepSeek-R1 demonstrated superior performance in supplementary informational value and readability, highlighting the importance of evaluating LLMs not only for accuracy but also for their capacity to enhance clinical communication and deci...

Li Xu, Xu Qiu, Jia-Yi Deng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Accuracy: Counterfactual Fragility and Demographic Bias in Clinical Evaluation of LLMs

Clinical LLM evaluation often emphasizes answer accuracy; however, accuracy alone does not test counterfactual consistency or demographic robustness. We evaluated six LLMs on 150 MedQA USMLE questions using two automated perturbation tests to assess their performance. The counterfactual validity (CFV) test asked each m...

Chaitai Deb Purkayastha, B. Bolla, Vishnu Surya Reddy Nandi · 0 citations
Review Open access Sep 2026

Severity Stratification Changes What Hallucination Rates Mean: An Item-Matched Audit of Five Large Language Models in Orthodontic Decision Support

Background/Objectives: Large language models (LLMs) are consulted for clinical decision support, and their reliability is judged by hallucination prevalence—a measure treating every unsupported element as equivalent, although a fabricated citation and a fabricated protocol differ in what a clinician acting on them woul...

Ferdi Allaf, Mustafa Özcan · 0 citations
Open access Aug 2026

Social Status and Clinical Resource Allocation by a Large Language Model: An Evaluation of 30,618 Decisions

Clinical use of LLM-based allocation support should therefore require explicit safeguards and systematic auditing for non-clinical influences, and their implicit and unexplained incorporation into resource-allocation decisions raises concerns regarding transparency, accountability, and clinical governance.

Siddharth Gandhi, Michael Balas · 0 citations
Open access Sep 2026

Psychometric validation and predictive efficacy of a comprehensive depression risk model for undergraduates

Background Undergraduate depression is prevalent, yet traditional screening is unidimensional and inefficient. We developed a biopsychosocial risk classification model for the cross-sectional identification of current depressive symptoms. Methods A cross-sectional study enrolled 898 undergraduates from a medical univer...

Xue Liang, Liu-Yi Lu, Qian Liao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.