Evaluating a large language model (ChatGPT-5) for detecting potential drug–drug interactions in intensive care: a cross-sectional comparative study with a clinical decision support system
Although ChatGPT-5 demonstrated limited diagnostic performance and the ability to generate clinically interpretable explanations, its low specificity and limited agreement with a rule-based system highlight important safety concerns, these findings suggest that LLMs may serve as complementary tools rather than standalo...