A pipeline for evaluating large language models (LLMs) on the task of capturing symptoms from clinical encounters and a gold standard dataset of symptom annotations from simulated doctor-patient encounter excerpts are developed.
Abstract
Abstract Objectives To develop a pipeline for evaluating large language models (LLMs) on the task of capturing symptoms from clinical encounters. Materials and Methods We created a gold standard dataset of symptom annotations from simulated doctor-patient encounter excerpts (264 encounters; 16 symptoms; double-coded and adjudicated). Nine different LLMs from 4 vendors (OpenAI, Meta, DeepSeek, Moonshot AI) were used as examples to test our evaluation pipeline; outputs were assessed for correct structure and symptom information. Results Of 3085 excerpts, 2087 (68%) contained symptoms. Pain, cough, and shortness of breath were most common; LLMs achieved F1 scores ranging 0.66-0.88 for these symptoms with minimal prompt engineering. Of tested models, GPT-4.1 demonstrated the best overall performance. Discussion Our evaluation pipeline and benchmarking dataset are publicly available and applicable to various LLMs, including open-source models. Conclusion This work supports the development and optimization of models that seek to improve patient symptom understanding.
LLM agents perform reliably for question generation and SAP drafting but require expert verification of formula composition, cohort boundary logic, and concordance computation before results are reported.
Yi-Lan Wu, D. J. Fu, Yu-Kun Zhou et al.· Journal of Medical Internet...· 0 citations
Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.
L. Barrett, N. Joshi, A. S. North et al.· medRxiv· 0 citations
The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.
Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al.· Cadernos de Saúde Pública· 1 citation
This article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs by describing underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question.
S. E. McKinney, P. Vu, S. Justice et al.· 0 citations
Proprietary LLMs, especially in few-shot (O1) or fine-tuned (Gemini 2.0 Pro) settings, significantly outperformed other models and confirm the power of modern LLMs for genomic knowledge extraction.
Claudiu Creanga, Teodor-George Marchitan, L. Dinu· International Conference on...· 0 citations
This entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.
Qiao Jin, Nicholas Wan, Robert Leaman et al.· Nature Protocols· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.