Evaluating the proficiency of ChatGPT-3.5, ChatGPT-4o and Google Gemini in responding to frequently asked questions on gender affirming care by transgender people
Abstract
Background: Transgender and gender-diverse individuals encounter stigma, discrimination, and healthcare access barriers, leading them to seek online information, including from large language models like ChatGPT and Google Gemini. This study compares their responses on gender-affirming care. Methods: This cross-sectional study evaluated responses from ChatGPT-3.5, ChatGPT-4o, and Google Gemini to 18 commonly asked questions on gender-affirming care, derived from current guidelines and clinical practice inquiries. Two endocrinologists independently assessed guideline adherence, while response reliability and quality were evaluated using the modified DISCERN scale and the Global Quality Scale, respectively. Hallucination tendency was rated using a Likert scale, and readability was analyzed using the Flesch Reading Ease, Flesch-Kincaid Grade Level, and Gunning Fog Index. Results: The highest guideline compatibility was observed with ChatGPT-4o [81.1%], followed by ChatGPT-3.5 [77.2%] and Google Gemini [63.8%]. For quality ChatGPT-4o scored highest [4.8±0.3], ChatGPT-3.5 slightly lower [4.5±0.5], and Gemini significantly lower [2.4±0.5]. ChatGPT-4o and ChatGPT-3.5 showed similar quality [p=0.19] but both outperformed Gemini [p<0.001]. ChatGPT-4o had the greatest tendency to hallucinate. Overall, a high school education was needed to understand the responses. In terms of readability, Gemini’s answers were easiest, followed by ChatGPT-3.5 and ChatGPT-4o [p<0.001]. Conclusions: ChatGPT-4o demonstrated the highest accuracy, quality, and guideline adherence in gender-affirming care, outperforming ChatGPT-3.5 and Google Gemini. However, its greater hallucination tendency necessitates caution in medical use. While Google Gemini had better readability, its lower compatibility and quality scores reduce its effectiveness in transgender care. These inaccuracies highlight that AI should not replace expert healthcare guidance.