Aug 2026· Journal of Personalized Medicine· Vol 16· 0 citations· 27 references
Medicine
TL;DR
Clinical use of LLM-based allocation support should therefore require explicit safeguards and systematic auditing for non-clinical influences, and their implicit and unexplained incorporation into resource-allocation decisions raises concerns regarding transparency, accountability, and clinical governance.
Abstract
Objective: The objective was to quantify whether demographic and social attributes that were irrelevant to stated clinical need, prognosis, and expected benefit altered resource-allocation decisions made by a general-purpose large language model (LLM). Methods: We conducted a cross-sectional audit of the gpt-5-chat-latest API model alias on 8 October 2025, across seven clinical vignettes, generating 30,618 forced-choice comparisons between patient profiles. Profiles varied across a full-factorial combination of eight demographic and social attributes while clinical need, prognosis, and expected benefit were held constant. Forced choices were analyzed using pooled logistic regression with separate Patient A and Patient B attribute terms and vignette-specific position effects; position-averaged odds ratios and position-balanced absolute probabilities were derived from this model. Priority-score differences were analyzed using an analogous linear model. Results: The model showed large position-averaged associations between non-clinical patient attributes and allocation decisions. Indigenous and Black race were associated with substantially higher odds of selection relative to White race (Indigenous: OR 16.48, 95% CI 14.85–18.28; Black: OR 8.07, 95% CI 7.32–8.90), corresponding to position-balanced absolute increases in selection probability of 30.7 and 16.3 percentage points, respectively. Conversely, high-status occupation (OR 0.064, 95% CI 0.058–0.071), friendship with institutional leadership (OR 0.121, 95% CI 0.111–0.131), and major donor status (OR 0.092, 95% CI 0.084–0.101) were associated with markedly lower odds of selection. Choice-score concordance was 95.1%. Conclusions: In this controlled audit, the LLM’s allocation decisions varied substantially according to demographic and social characteristics despite identical stated clinical need, prognosis, and expected benefit. Although some patterns could be interpreted differently under competing ethical frameworks, their implicit and unexplained incorporation into resource-allocation decisions raises concerns regarding transparency, accountability, and clinical governance. Clinical use of LLM-based allocation support should therefore require explicit safeguards and systematic auditing for non-clinical influences.
It is shown that requiring sufficient individual-level SDoH survey data results in significant selection bias and sample reduction in AoU, and that area-level SDoH metrics contribute to disease prediction independently of individual-level measures.
M. Hysong, A. Manning, Michael D. Green et al.· Communications Health· 1 citation
The three models answered many late-life depression questions accurately and safely, but none demonstrated consistently reliable performance across complex or high-risk scenarios.
Wei Xiao, Huan Zhang, Xiao-Yi Chen et al.· Frontiers in Psychiatry· 0 citations
HIVMedQA is developed, a clinician-curated benchmark of HIV-related open-ended medical question-answer pairs spanning basic knowledge, clinical reasoning, complex patient vignettes, and bias-modified scenarios that provides a structured benchmark for evaluating LLMs in HIV clinical decision support.
Gonzalo Cardenal-Antolin, J. Fellay, Bashkim Jaha et al.· Communications Medicine· 1 citation
LLM clinical recommendations varied by sociodemographic descriptors, particularly under diagnostic uncertainty, suggesting structured inputs may mitigate bias in clinical AI systems.
Idit Tessler, Mahmud Omar, A. Wolfovitz et al.· Journal of Vestibular Resear...· 0 citations
Publicly accessible English-language web-interface outputs from current LLMs showed systematic demographic patterns in pediatric obesity risk attribution, supporting the need for pre-deployment and post-deployment bias auditing before clinical or consumer health use.
Can Wang, Zhen-Dong Liu, Yan-Yu Jiang et al.· Frontiers in Public Health· 0 citations
LLM-derived mobility functional status assessment was associated with ED visits and hospitalization, but not with all-cause mortality after adjustment, and prospective validation and evaluation of incremental discrimination, calibration, and clinical utility are needed before clinical implementation.
Sandeep R. Pagali, Xing-Yi Liu, He-Ling Jia et al.· Journal of The American Geri...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.