Open access
Jul 2026
Beyond accuracy: evaluating the reliability of large language models for medical assessment
For automated assessment metadata extraction, reliability rather than accuracy determines whether an LLM can be deployed, and national origin is not a meaningful predictor.
Hui Zhang, Lihui Qu, Hongbo Bai et al.
· Frontiers in Artificial Inte... · 0 citations