Review
Open access
Aug 2026
The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review.
It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.
Euijun Yang, S. Ko, Hyekyung Woo
· Journal of Medical Internet... · 0 citations