Benchmarking five large language models in medical genetics: a bilingual comparative evaluation using published and novel expert-authored questions
This study asked whether five contemporary large language models answer medical genetics multiple-choice questions with equivalent accuracy on published versus novel items and across English and Turkish, and sought to characterize the errors that persist. Five models (GPT-5.2, Gemini 3 Pro, Claude Sonnet 4.6, Grok 4,...