Skip to content
Open access

When Artificial Intelligence Speaks For The Obstetrician: Multilingual Accuracy On Real Patient Questions

Aug 2026 · Hippocrates Medical Journal · 0 citations · 24 references

TL;DR

This study aimed to compare the response accuracy and reference quality of three free LLMs in Turkish and English using common pregnancy-related questions and found that Google Gemini and DeepSeek provided more accurate responses in English than in Turkish.

Abstract

BackgroundThe use of generative large language models (LLMs) in healthcare is rapidly increasing, offering easier access to medical information. However, comprehensive data on their multilingual accuracy and the reliability of cited scientific references remain limited. This study aimed to compare the response accuracy and reference quality of three free LLMs (ChatGPT, Google Gemini, DeepSeek) in Turkish and English using common pregnancy-related questions.MethodsIn this comparative observational study, 14 frequently asked pregnancy questions were posed to each LLM in Turkish and English, requesting responses supported by up-to-date scientific web sources. Answers were evaluated blindly by obstetricians and gynaecologists for accuracy. References were independently assessed for reliability, scientific validity, and accessibility. Statistical analyses were performed.ResultsLanguage and model infrastructure significantly influenced performance. Google Gemini and DeepSeek provided more accurate responses in English than in Turkish (p

Read PDF

Similar papers

Sep 2026

Evaluation of the accuracy and reproducibility of large language models (ChatGPT, DeepSeek, Gemini) in responding to patient-centered lipedema questions.

BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered informatio...

Rabia Sanır, E. Türkmen, E. Giray et al. · 0 citations
Review Open access Sep 2026

Assessing the accuracy and usability of artificial intelligence-based language models in responding to common periodontal patient questions.

BACKGROUND Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aim...

Sajad Jahantigh, R. Amid, A. Moscowchi et al. · 0 citations
Review Open access Aug 2026

Accuracy and Readability of Generative Artificial Intelligence for Vascular Surgery Patients: A Specialist Based Evaluation Highlighting the Current Landscape of Safety Risks and Accessibility Gaps.

OBJECTIVE Patients frequently seek health information and medical advice from chatbots instead of consulting their physicians or referring to credible patient education resources provided by medical societies. A study to evaluate the quality, readability, and clinical appropriateness of ChatGPT generated answers to com...

Mario D'Oria, W. Dorigo, V. Alexiou et al. · 0 citations
#natural language process... Preprint Sep 2026

HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication

Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication. Fo...

Hassan Saeed Hassan Albattra, Mazen Bahgat, Rahatara Ferdousi et al. · 0 citations
Sep 2026

Comparative evaluation of large language models in interpreting the scientific literature on intraoral scanners across varying input levels.

STATEMENT OF PROBLEM Dental professionals increasingly seek artificial intelligence (AI)-generated sources for concise information on emerging topics in advancement and scientific literature; however, concerns persist regarding the accuracy of AI tools and the reliability of responses generated based on the complexity...

Pranay Jain, F. R, A. V et al. · 0 citations
Open access Sep 2026

Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer

Background Breast cancer, the most common malignancy among women, remains a major global public health concern. With the rapid growth of artificial intelligence–based language models, it is essential to evaluate their potential roles in patient education. This study compared ChatGPT and Gemini in responding to patient-...

Yeliz Yılmaz Bozok, N. Acar, M. Atahan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.