Open access
Jul 2026
Development of a benchmarking dataset for symptom detection using large language models
A pipeline for evaluating large language models (LLMs) on the task of capturing symptoms from clinical encounters and a gold standard dataset of symptom annotations from simulated doctor-patient encounter excerpts are developed.
Joshua Davis, B. Durieux, Chloe Van Dongen et al.
· JAMIA Open · 0 citations