Skip to content
Review Open access

Development and Clinical Validation of an Automated LLM Judge for Evaluating Perioperative Patient Questions

Sep 2026 · Bioengineering · 0 citations · 28 references

Abstract

Background: Large language models (LLMs) are increasingly used in patient-facing clinical applications, creating a need for scalable methods to evaluate the accuracy, safety, and appropriateness of their responses. Although physician review remains the reference standard, it is resource-intensive and difficult to scale. Objective: To develop and clinically validate an LLM-as-a-judge framework for evaluating responses to patients’ periprocedural questions. Methods: We retrospectively evaluated 1285 patient–AIVA interactions from 112 patients across two Mayo Clinic sites and multiple surgical specialties. Physician reviewers classified interactions using a predefined true-positive (TP), false-negative (FN), true-negative (TN), and false-positive (FP) framework. The finalized LLM judge independently evaluated the same interactions. Agreement was assessed using four-class and category-specific agreement, Cohen’s kappa, and a secondary binary analysis of response correctness. Results: Overall, four-class agreement was 93.1% (95% CI, 91.7–94.5%), with Cohen’s κ = 0.852 (95% CI, 0.823–0.882). Category-specific agreement was 94.1% for TP, 90.7% for FN, 95.5% for TN, and 84.9% for FP classifications. In the binary correctness analysis, accuracy was 93.4% (95% CI, 92.0–94.7%), sensitivity 94.6%, specificity 90.2%, precision 96.2%, and F1-score 0.954. McNemar’s test demonstrated no significant asymmetry between paired classifications (p = 0.11). Conclusions: A clinically grounded LLM judge closely approximated physician evaluation of patient-facing AI responses. These findings support its potential as a scalable assistive monitoring tool while preserving physician oversight for uncertain or safety-sensitive interactions.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.