Skip to content

Benchmarking large language models for multilingual health information delivery: a mixed-methods data analytics framework applied to vaccine communication in a middle-income setting

Aug 2026 · International Journal of Data Science and Analysis · Vol 22 · 0 citations · 37 references

TL;DR

This framework provides a replicable approach to evaluating AI-generated health content across languages and contexts and identified five influential factors: local contextualisation and regulatory anchoring, structured presentation and readability, balance between completeness and accessibility, source credibility, and culturally sensitive tone.

View source

Similar papers

Review Open access Sep 2026

Benchmarking five large language models in medical genetics: a bilingual comparative evaluation using published and novel expert-authored questions

This study asked whether five contemporary large language models answer medical genetics multiple-choice questions with equivalent accuracy on published versus novel items and across English and Turkish, and sought to characterize the errors that persist. Five models (GPT-5.2, Gemini 3 Pro, Claude Sonnet 4.6, Grok 4,...

Özge Beyza Gündoğdu Öğütlü, Benjamin D. Solomon, Y. Çelik · 0 citations
Open access Aug 2026

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next, and findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Nıyazı Çetın, A. Atılan · 0 citations
Open access Aug 2026

Evaluating search-enabled large language model interfaces for mpox public health consultation: a guideline-based comparative study

The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability, which support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are eva...

Qi-Qi Zheng, Ru Chen, Ming-Ming Cai et al. · 0 citations
Open access Sep 2026

Context Matters in LLM-Assisted Qualitative Data Analysis: Workflow Development and Multidimensional Evaluation in Health Research

Large language models (LLMs) are increasingly used for qualitative analysis, but strong language performance does not guarantee strong interpretation when meaning depends on method and context. We developed a context-specific workflow that translated framework analysis into bounded, sequential LLM-supported tasks and a...

H. Dai, Y. Li, M. Villalobos-Quesada et al. · 0 citations
Open access Aug 2026

Evaluating large language models using the Type 2 Diabetes Health Education guideline: a comparative analysis of ChatGPT-4.1, Claude-4.0, DeepSeek-V3, and ERNIE Bot 4.5 Turbo

Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education.

Zhaoxia Huang, Yuxin Bai, Jun-Yue Luo et al. · 0 citations
Review Open access Sep 2026

Evaluating large language models as clinical decision support tools in primary healthcare settings: Protocol for a multi-country comparative validation study on expert-adjudicated hypothetical vignettes (hypMOOVE-PHC)

The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania, and aims to validate a pool of LLMs through clinical review of expert-generated vignettes through fully crossed repeated-measures comparative eva...

P. Macharia, C. Kachimanga, M. Mahende et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.