Benchmarking large language models for multilingual health information delivery: a mixed-methods data analytics framework applied to vaccine communication in a middle-income setting
Aug 2026· International Journal of Data Science and Analysis· Vol 22· 0 citations· 37 references
TL;DR
This framework provides a replicable approach to evaluating AI-generated health content across languages and contexts and identified five influential factors: local contextualisation and regulatory anchoring, structured presentation and readability, balance between completeness and accessibility, source credibility, and culturally sensitive tone.
This study asked whether five contemporary large language models answer medical genetics multiple-choice questions with equivalent accuracy on published versus novel items and across English and Turkish, and sought to characterize the errors that persist. Five models (GPT-5.2, Gemini 3 Pro, Claude Sonnet 4.6, Grok 4,...
Özge Beyza Gündoğdu Öğütlü, Benjamin D. Solomon, Y. Çelik· Frontiers in Medicine· 0 citations
How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next, and findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.
Nıyazı Çetın, A. Atılan· European Journal of Therapeu...· 0 citations
The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability, which support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are eva...
Qi-Qi Zheng, Ru Chen, Ming-Ming Cai et al.· Frontiers in Public Health· 0 citations
Large language models (LLMs) are increasingly used for qualitative analysis, but strong language performance does not guarantee strong interpretation when meaning depends on method and context. We developed a context-specific workflow that translated framework analysis into bounded, sequential LLM-supported tasks and a...
H. Dai, Y. Li, M. Villalobos-Quesada et al.· medRxiv· 0 citations
Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education.
Zhaoxia Huang, Yuxin Bai, Jun-Yue Luo et al.· Frontiers in Public Health· 0 citations
The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania, and aims to validate a pool of LLMs through clinical review of expert-generated vignettes through fully crossed repeated-measures comparative eva...
P. Macharia, C. Kachimanga, M. Mahende et al.· medRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.