Bottom-up incremental scoring showed the closest alignment with human assessment in clinical AI evaluation, underscoring the need for standardised prompt architectures in clinical AI evaluation.
Abstract
INTRODUCTION
Large language model (LLM)-as-a-judge systems offer scalable evaluation of artificial intelligence (AI)-generated clinical outputs, yet their susceptibility to prompt variability raises concerns regarding reproducibility and alignment with expert judgement. This study examined whether evaluation prompt strategies influence scoring patterns and concordance with clinical raters in critical care.
MATERIAL AND
Methods
This post-hoc analysis used 90 structured clinical reports generated in a prior study using an XGBoost ICU mortality prediction model trained on the MIMIC-IV database. GPT-4o (Azure AI, version 2024-11-20) produced structured interpretations from risk estimates and SHAP attributions. These outputs were evaluated using the IMPACT framework under three evaluation prompt strategies: baseline (E1), top-down decremental (E2), and bottom-up incremental (E3). Agreement between clinician ratings and the automated o3-mini evaluator (Azure AI, version 2025-01-31) was assessed using intraclass correlation coefficients (ICC), with strategy comparisons by Fisher's z-transformation. Score deviations were examined with repeated-measures ANOVA.
Results
Mean IMPACT scores were 79.9 (SD 9.9) for E1, 83.3 (SD 9.6) for E2, 78.7 (SD 9.1) for E3, and 78.6 (SD 8.9) for clinicians. All strategies demonstrated substantial agreement (ICC > 0.80). E2 showed significantly lower agreement with clinicians (ICC = 0.82) than E1 and E3 (both ICC = 0.94, p < 0.001). Score deviations differed significantly across strategies (p < 0.001), with E3 showing the smallest mean deviation (0.1) and E2 the largest (4.7).
Conclusions
Prompt design meaningfully affects both IMPACT scoring patterns and the reliability of LLM-based evaluators. Bottom-up incremental scoring showed the closest alignment with human assessment, underscoring the need for standardised prompt architectures in clinical AI evaluation.
Evaluating generative AI output remains a critical bottleneck for safe and scalable deployment of AI in healthcare. Expert clinical judgement is often presented as the gold standard, but human assessment is costly and inconsistent. LLM-as-judge systems, i.e., leveraging AI to evaluate other AI outputs, have been proposed, yet their reliability in global health remains untested. We compared five LLM judges and six human clinicians in evaluating responses to questions posed by Rwandan health workers. The highest-performing LLM-judge (Claude-4.1-Opus) matched human evaluators on only four of eleven evaluation criteria, with other models scoring too leniently (Gemini-2.5-Pro) or too harshly (GPT-5). Constructing LLM-juries to balance model-specific biases improved agreement on only one additional criterion. Notably, performance and cost-effectiveness fell when moving from English to Kinyarwanda. Overall, while LLM-judges show promise, their inability to handle linguistic and cultural context is a critical limitation, underscoring the need for further investment in scalable evaluation solutions.
G. Williams, S. Rutunda, Floris Nzabakira et al.· npj Digital Medicine· 1 citation
Clinical documentation in Electronic Health Records (EHRs) remains a substantial source of administrative burden for clinicians. In this study, we evaluate a modular AI-assisted clinical documentation pipeline using two complementary approaches: (1) a controlled benchmark based on multilingual synthetic clinical dialogues, and (2) an observational analysis of real-world usage traces from routine deployments. The benchmark enables systematic comparison of ASR–LLM configurations under fully controlled conditions, using metrics for transcription accuracy (Word Error Rate and Medical WER), report-generation quality, and modeled processing cost. Within this benchmark setting, Voxtral showed the strongest ASR performance among the evaluated models, while GPT-4o and Gemini 1.5 Pro showed the strongest report-generation performance under the automated evaluation used in this study. The real-world trace analysis should be interpreted as descriptive evidence of operational use, not as prospective clinical validation or as a direct evaluation of any single benchmarked configuration. Taken together, the results support the use of this pipeline as a human-supervised draft-generation tool that still requires clinician review, local workflow evaluation, and prospective clinical validation before broader deployment.
Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured"safety gain"reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3) answered a fully paired common panel (1,200 cells) with a standard prompt and the wrapper. The pre-specified endpoint was the paired reduction in unsafe overconfidence scored by the primary judge (GPT-5.4-nano); secondary analyses added a different-family judge (Claude Sonnet 5), a correctness judge, matched scaffold controls, and a blinded three-clinician review. Results: Unsafe overconfidence fell from 49.3% to 24.7%, a paired reduction of 24.7 points (95% CI 21.8-27.7; p<0.001), robust in direction across models and paraphrases. Magnitude was judge-dependent: Sonnet agreed on direction but nearly halved the effect (+13.1 points), with one-directional disagreement. Blinded clinicians characterized the primary judge as a high-sensitivity (1.00), low-specificity (0.55) screen, not a calibrated rate. The gain carried a model-specific helpfulness cost (correct diagnosis 80.3% to 50.3%): near-free for GPT-5.5, near-total for Gemini (-58 points). Matched scaffold controls showed genuine behavior change, not judge circularity. Conclusions: LLM-judged clinical safety effects should be reported as directional and relative, anchored to human review and evaluated jointly with helpfulness, not as calibrated absolute rates. This does not establish clinical deployment readiness.
A single LLM verifier lacks sufficient reliability to serve as a stand-alone judge of clinical reasoning at scale, and that structured human oversight remains essential.
Hyunjung Byun, Dahyoun Lee, Munyoung Jung et al.· Journal of medical systems· 1 citation· ⚡1
The reliability and robustness of LaaJ for specialized medical knowledge is investigated by evaluating six LLMs for their judgment capabilities on three dimensions: correctness, readability, and completeness, and it is observed that hallucinations in LaaJ setups can be mitigated by epistemic markers.
Ioana Buhnila, Aman Sinha, Rohit Agarwal et al.· BioNLP@ACL· 0 citations
This study investigates the readability, clinical reliability, and temporal consistency of artificial intelligence (AI) chatbots regarding pneumothorax information. A question bank comprising 40 patient-centered queries was deployed across three large language models (ChatGPT, Gemini, Copilot), stratified by two access tiers and two prompting strategies (zero-shot versus the optimized PROMPORT strategy). Queries were replicated longitudinally on Days 1, 3, and 7 under strict session-control protocols. Text accessibility was quantified using five automated readability indices, while two independent, blinded thoracic surgeons evaluated clinical quality using modified DISCERN (mDISCERN), JAMA benchmarks, and PEMAT-P indices. Readability metrics demonstrated absolute structural stability across the tracking intervals (p > 0.05). Unprompted configurations consistently generated complex, high-school-level outputs, whereas the PROMPORT strategy successfully compressed linguistic variances and neutralized chronological algorithmic drift (p > 0.05). Conversely, unprompted architectures exhibited significant temporal volatility in mDISCERN and JAMA profiles (p < 0.05), which was successfully stabilized by optimized prompt constraints. Inter-rater reliability was high across all structural evaluations. In conclusion, while unprompted models exhibit marked baseline linguistic and quality variations, the strategic integration of robust prompt engineering successfully enforces the temporal stability and clarity required for reliable digital public health communication.
Ömer Önal, Suzan Temiz Bekce· Scientific Reports· 0 citations