Safety-Oriented Benchmarking of Large Language Models in Risk-Based Management of Abnormal Cervical Screening Results: Scenario-Based Benchmark Study.
BACKGROUND Large language models (LLMs) are increasingly being considered for clinical decision support, yet their safety in risk-based cervical screening management remains insufficiently characterized. OBJECTIVE This study benchmarked the guideline concordance and safety-related performance of 3 LLMs in the initial...