Skip to content
Open access

Poster 162. Racial Biases Perpetuated by Modern Large Language Models Negatively Impact Diagnostic Reasoning and Treatment Recommendations in Musculoskeletal Healthcare

Aug 2026 · Orthopaedic Journal of Sports Medicine · Vol 14 · 0 citations

Abstract

Objectives: To determine whether contemporary large language models (LLMs) perpetuate gender and racial biases in medical decision-making. Methods: A total of 180 standardized vignettes concerning musculoskeletal diagnoses and treatments were extracted from AAOS Restudy and Orthobullets. A customized GPT-o4-mini model was prompted using designs that resembled typical use within clinical settings and fed vignettes via a custom queryLLM2 pipeline, programmatically generating JSON-formatted differentials across demographic groups (Caucasian/African-American/Asian/Hispanic and Male/Female) while holding all other content constant. JSON outputs were parsed to extract ordered diagnosis lists, compute the rank of the reference ('correct') diagnosis and a top-3 inclusion indicator, and derive sentiment scores. Differential diagnosis and treatment planning were evaluated across demographic groups using Kruskal-Wallis tests and pairwise Mann-Whitney U tests with Benjamini-Hochberg FDR correction, alongside standardized proportion and rank-delta visualizations. Results: Among 1,440 race/gender combinations, the model demonstrated outputs that were more likely to recommend diagnoses and treatments that stereotyped certain racial groups (Figure 1). Furthermore, the model was significantly more likely to provide the correct diagnosis for Asian and Caucasian patients, while the proportion of correct diagnoses within the top 3 diagnoses listed was significantly lower for African American and Hispanic patients (56% vs. 27%, p<0.05). Gender was not significantly associated with different diagnostic or treatment rankings. Conclusions: Contemporary LLMs may perpetuate racial biases acquired during model training when being used to reason through musculoskeletal healthcare content. These findings highlight a concerning limitation in the use of LLMs and therefore there is a need for enhanced transparency and mitigation of these biases prior to integration into clinical workflows.

Read PDF

Similar papers

Aug 2026

Can Large Language Models Preserve Diagnostic Accuracy Despite Patient Self-Diagnosis and Framing Bias in Hypothetical Upper-Extremity Scenarios?

PURPOSE Online health queries are often addressed by large language models (LLMs) embedded in search engines. It is possible that LLMs, like human clinicians, might be misdirected by vague symptom descriptions or inaccurate self-diagnoses. We examined patient and scenario factors associated with an LLM's ability to ide...

Emily H. Jaarsma, David C. Ring, John Wickman et al. · 0 citations
#large language models Open access Sep 2026

GPT-4 improves sex-specificity in cardiovascular patient education but may perpetuate gender biases: A mixed-methods audit

Patient education materials for cardiovascular disease (CVD) prevention frequently omit clinically important differences in disease manifestation, risk, and prevention between sexes. Socially constructed gender norms further shape how health information is communicated and received. Large Language Models (LLMs) like GP...

Samah Khan, G. Vaidean · 0 citations
Sep 2026

Can AI Predict Publication? Multimodal Large Language Models and the Structural Determinants of Surgical Scholarship.

BackgroundWhether artificial intelligence can identify publishable scientific work is untested. We evaluated whether a multimodal large language model (MLLM) could predict, from poster content alone, which abstracts at the American Association for the Surgery of Trauma (AAST) Annual Meetings reached publication, and ch...

Sohail Khan, Gavin McAfee, Alex Chiodo Ortiz et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment

Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future information rather than reasoning under the uncertainty present at the decision point. We introduc...

Misaki Matsuura, Sayantan Kumar, Ojas Kadam et al. · 0 citations
Open access Sep 2026

Sex and gender bias in large language models: an old problem at a new scale.

Patients are increasingly turning to large language models for health information, which makes the consistency of these systems' outputs across patient groups a public health concern. Demographic bias is harder to see than fabrication because it escapes accuracy benchmarks and shows up instead in the language the model...

C. Barbati, V. Casigliani, C. Rizzo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.