Aug 2026· Journal of Vestibular Research-Equilibrium & Orientation· pp.
9574271261474218
· 0 citations· 16 references
Medicine
TL;DR
LLM clinical recommendations varied by sociodemographic descriptors, particularly under diagnostic uncertainty, suggesting structured inputs may mitigate bias in clinical AI systems.
Abstract
ObjectiveAs large language models (LLMs) enter clinical decision support, concerns persist about sociodemographic bias. We assessed whether LLM recommendations for dizziness vary by patient descriptors and clinical detail.MethodsWe conducted a cross-randomized in-silico vignette study. One hundred synthetic emergency department dizziness cases were created using established diagnostic frameworks including the TiTrATE paradigm, SAEM GRACE-3 guidelines, and Bárány Society diagnostic criteria. Each vignette was tested in a neutral form and with 33 sociodemographic descriptor variants (34 total). Twelve instruction-tuned LLMs from multiple model families were evaluated. Models answered five binary clinical decision questions addressing etiology classification, triage disposition, neuroimaging, bedside vestibular examination, and mental health referral. Each model-vignette-descriptor combination was repeated 10 times, yielding 2,040,000 responses. Sociodemographic bias was quantified as descriptor-specific percentage-point deviations from neutral control recommendations with 95% confidence intervals.ResultsSociodemographic descriptors influenced LLM recommendations, with the largest differences observed for mental health referral decisions in diagnostically ambiguous cases. Referral likelihood was lower for Black transgender women (-12.2 pp; 95% CI -14.0 to -10.3), Black patients experiencing homelessness (-9.1 pp; -11.0 to -7.3), and patients experiencing homelessness (-7.7 pp; -9.5 to -5.9). Differences were attenuated when vignettes contained clearer diagnostic information. Other effects were smaller, including increased neuroimaging recommendations for low-income descriptors (+4.0 pp; 95% CI 2.1-5.8).ConclusionLLM clinical recommendations varied by sociodemographic descriptors, particularly under diagnostic uncertainty. More detailed clinical information reduced these disparities, suggesting structured inputs may mitigate bias in clinical AI systems.
While all four models exhibited comparable adherence to guideline-based content, DeepSeek-R1 demonstrated superior performance in supplementary informational value and readability, highlighting the importance of evaluating LLMs not only for accuracy but also for their capacity to enhance clinical communication and deci...
Li Xu, Xu Qiu, Jia-Yi Deng et al.· Neurology Asia· 0 citations
Clinical LLM evaluation often emphasizes answer accuracy; however, accuracy alone does not test counterfactual consistency or demographic robustness. We evaluated six LLMs on 150 MedQA USMLE questions using two automated perturbation tests to assess their performance. The counterfactual validity (CFV) test asked each m...
Chaitai Deb Purkayastha, B. Bolla, Vishnu Surya Reddy Nandi· 0 citations
Background/Objectives: Large language models (LLMs) are consulted for clinical decision support, and their reliability is judged by hallucination prevalence—a measure treating every unsupported element as equivalent, although a fabricated citation and a fabricated protocol differ in what a clinician acting on them woul...
Ferdi Allaf, Mustafa Özcan· Diagnostics· 0 citations
Clinical use of LLM-based allocation support should therefore require explicit safeguards and systematic auditing for non-clinical influences, and their implicit and unexplained incorporation into resource-allocation decisions raises concerns regarding transparency, accountability, and clinical governance.
Siddharth Gandhi, Michael Balas· Journal of Personalized Medi...· 0 citations
Background Undergraduate depression is prevalent, yet traditional screening is unidimensional and inefficient. We developed a biopsychosocial risk classification model for the cross-sectional identification of current depressive symptoms. Methods A cross-sectional study enrolled 898 undergraduates from a medical univer...
Xue Liang, Liu-Yi Lu, Qian Liao et al.· Frontiers in Psychiatry· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.