Beyond accuracy: an educational benchmarking study of task fragility and reasoning stability of large language models on dermatology board-style questions.
BACKGROUND Large language models (LLMs) are increasingly used in medical education and assessment frameworks. However, their reliability under varying task demands remains unclear. Significant gaps exist in the literature regarding the reasoning stability of these models when faced with fluctuations in question structu...