A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology (RadM-Bench): Evaluation Development and Validation
Abstract Background Multimodal large language models are increasingly used in radiological diagnosis, but their performance has not been systematically evaluated across volumetric (3D) imaging, real-world clinical versus public teaching cases, and bilingual contexts. Objective The aim of the study is to develop a bilin...