Open access
Jul 2026
Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases.
While AI agents delivered highly efficient, directionally aligned assessments, they did not fully capture the nuances of human clinical judgment and could not substitute for physician-centered evaluation and promise assistive tools that can triage or pre-screen outputs to reduce human burden.
Peilun Shi, Jian Li, Ziqi Yang et al.
· npj Digital Medicine · 0 citations