Aug 2026· Hippocrates Medical Journal· 0 citations· 24 references
TL;DR
This study aimed to compare the response accuracy and reference quality of three free LLMs in Turkish and English using common pregnancy-related questions and found that Google Gemini and DeepSeek provided more accurate responses in English than in Turkish.
Abstract
BackgroundThe use of generative large language models (LLMs) in healthcare is rapidly increasing, offering easier access to medical information. However, comprehensive data on their multilingual accuracy and the reliability of cited scientific references remain limited. This study aimed to compare the response accuracy and reference quality of three free LLMs (ChatGPT, Google Gemini, DeepSeek) in Turkish and English using common pregnancy-related questions.MethodsIn this comparative observational study, 14 frequently asked pregnancy questions were posed to each LLM in Turkish and English, requesting responses supported by up-to-date scientific web sources. Answers were evaluated blindly by obstetricians and gynaecologists for accuracy. References were independently assessed for reliability, scientific validity, and accessibility. Statistical analyses were performed.ResultsLanguage and model infrastructure significantly influenced performance. Google Gemini and DeepSeek provided more accurate responses in English than in Turkish (p
BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered informatio...
Rabia Sanır, E. Türkmen, E. Giray et al.· Phlebology· 0 citations
BACKGROUND
Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aim...
Sajad Jahantigh, R. Amid, A. Moscowchi et al.· Clinical Advances in Periodo...· 0 citations
OBJECTIVE
Patients frequently seek health information and medical advice from chatbots instead of consulting their physicians or referring to credible patient education resources provided by medical societies. A study to evaluate the quality, readability, and clinical appropriateness of ChatGPT generated answers to com...
Mario D'Oria, W. Dorigo, V. Alexiou et al.· European Journal of Vascular...· 0 citations
Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication. Fo...
STATEMENT OF PROBLEM
Dental professionals increasingly seek artificial intelligence (AI)-generated sources for concise information on emerging topics in advancement and scientific literature; however, concerns persist regarding the accuracy of AI tools and the reliability of responses generated based on the complexity...
Pranay Jain, F. R, A. V et al.· The Journal of prosthetic de...· 0 citations
Background Breast cancer, the most common malignancy among women, remains a major global public health concern. With the rapid growth of artificial intelligence–based language models, it is essential to evaluate their potential roles in patient education. This study compared ChatGPT and Gemini in responding to patient-...
Yeliz Yılmaz Bozok, N. Acar, M. Atahan et al.· PLoS ONE· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.