Development and Clinical Validation of an Automated LLM Judge for Evaluating Perioperative Patient Questions
Background: Large language models (LLMs) are increasingly used in patient-facing clinical applications, creating a need for scalable methods to evaluate the accuracy, safety, and appropriateness of their responses. Although physician review remains the reference standard, it is resource-intensive and difficult to scale...