NCCN-anchored RAG outperformed both baseline GPT-5 and a literature-anchored clinical AI without direct guideline access, and external validation of guideline anchoring's operational importance provides external validation of guideline anchoring's operational importance.
Abstract
INTRODUCTION
Large language models (LLMs) are being studied as oncology decision-support tools but can produce inaccurate outputs. We compared LLM performance in gynecologic oncology across three knowledge-integration configurations differing in retrieval strategy and underlying model, using the modified Generative Performance Score (mGPS) as the primary outcome.
Methods
Fifty de-identified gynecologic oncology cases were submitted (October-November 2025) to three LLMs: baseline GPT-5, an NCCN-anchored GPT-5 retrieval-augmented generation (RAG) configuration, and OpenEvidence (a literature-anchored clinical AI without NCCN access at that time). Three gynecologic oncologists independently scored outputs using the mGPS (range - 1 to +1; Guideline Concordance plus Hallucination Penalty). Wilcoxon signed-rank tests and mixed-effects ordered logistic regression were used.
Results
GPT-RAG produced the highest mGPS (0.83, SD 0.26), followed by OpenEvidence (0.70, SD 0.27) and baseline GPT-5 (0.65, SD 0.31). GPT-RAG exceeded baseline (W = 189.5, Z = -3.42, P < .001, r = 0.49) and OpenEvidence (W = 254.0, Z = -2.64, P = .008); OpenEvidence and baseline did not differ (P = .22). Mixed-effects modeling confirmed higher mGPS for GPT-RAG (OR 3.74; 95% CI, 1.57-8.90). Inter-rater agreement (ICC) was 0.49 for mGPS, 0.30 for Hallucination Penalty, and 0.70 for Readability and Rationality.
Conclusion
NCCN-anchored RAG outperformed both baseline GPT-5 and a literature-anchored clinical AI without direct guideline access. OpenEvidence's subsequent NCCN integration (April 27, 2026) provides external validation of guideline anchoring's operational importance. Findings reflect benchmark performance, not clinical safety or improved patient outcomes.
Background/Objectives: Large language models (LLMs) show promise for clinical decision support, yet their accuracy in interpreting specialized medical guidelines remains uncertain. Retrieval-augmented generation (RAG) may enhance performance by grounding responses in authoritative knowledge bases. This study aimed to compare the accuracy, comprehensiveness, and safety of RAG-enhanced versus standard LLMs for answering clinical questions derived from the German S3 guideline for oral cavity carcinoma. Methods: We conducted a prospective, single-blind benchmark study evaluating six LLMs: one RAG-enhanced model (Custom GPT with guideline access), one consensus-based model (ConsensusGPT), and four standard models (DeepSeek-V3.2, Mistral Small 3.2, Qwen3-Next-80B, GPT-OSS-120B). Fifty clinical questions covering 17 guideline domains were presented to each model three times, yielding 900 evaluations. Three expert reviewers assessed responses using 5-point Likert scales for accuracy, comprehensiveness, and clarity, under a single-blind procedure, the effectiveness of which was tested by a pre-specified manipulation check. We then ran a paired within-model experiment in which each base model was queried with and without guideline access through a transparent, openly released retrieval pipeline, and scored every response with a condition-blind automated judge alongside deterministic retrieval metrics computed from the logs. Secondary outcomes included hallucination rates and guideline citation behavior. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs). Results: In a paired within-model design that held each base model fixed, adding transparent guideline retrieval improved accuracy—significantly in the three weaker open-weight models (Mistral, Qwen3, and GPT-OSS) and directionally in the already-strong DeepSeek and GPT-5 bases. Because a pre-specified blinding check found that experts could still identify retrieval-augmented answers with 98.5% accuracy, we anchored causal interpretation on measures that do not depend on the human raters, ranked by their independence: deterministic, log-derived retrieval metrics first, and then an automated, condition-blind LLM judge, whose agreement with the experts (Spearman ρ = 0.81, 95.7% within-one agreement) establishes shared calibration rather than independence from their bias. Deterministically from the retrieval logs, citation groundedness rose from 0% to 51–89% and retrieval recall@5 was 92%. On the judge, content-level hallucination fell from 42% to 4% and accuracy rose by a pooled +0.64 points (95% CI 0.47–0.80); the accuracy gain persisted after adjustment for response length (+0.48, 95% CI 0.22–0.73), which retrieval shortened rather than lengthened. The accuracy gain was large for weaker base models and small or non-significant for already-strong ones, whereas the hallucination and auditability gains were consistent across all models. The human ratings reproduced the judge’s accuracy effect (+0.61, 95% CI 0.49–0.74), and GPT-5 run through the transparent pipeline showed no significant difference from the proprietary Custom GPT (judge accuracy 4.48 vs. 4.58). Conclusions: Guideline retrieval yields a reproducible, largely base-independent improvement in the safety and auditability of LLM answers to clinical guideline questions, with accuracy gains concentrated in weaker base models. Because retrieval-augmented answers are recognizable to experts, rigorous evaluation should rely on rater-independent measures, and residual hallucination continues to require human oversight.
Andreas Vollmer, Lara Schorn, Felix Schrader et al.· Diagnostics· 0 citations
Background Radiology impressions guide clinical care. Large Language Models (LLMs)-drafted impressions can drift into generic, off-style text. Retrieval-augmented generation (RAG) enables context-aware few-shot prompting during inference. Methods This retrospective IRB-approved study included 11,998 CT pulmonary angiography (CTPA) reports. We built a retrieval bank from 11,399 reports and reserved 599 reports for testing. GPT-4o and LLaMA 3.1-70B generated impressions from the “findings” section using three setups: zero-shot, fixed random few-shot, and dynamic retrieval-selected few-shot (top-k semantic matches; k = 3/5/10). We ran temperatures 0, 0.7, 1. We scored outputs against the original impressions with ROUGE and BERTScore F1, report mean scores with 95% confidence intervals, and tested for statistical significance using Wilcoxon signed-rank test. Results Dynamic retrieval-based few-shot prompting outperformed zero-shot and fixed few-shot prompting across all configurations (all p < 0.05). The highest scores were observed at temperature 0 and k = 10. ROUGE-1 F1 increased to 0.44–0.47 for GPT-4o and 0.37–0.50 for LLaMA, versus 0.35–0.37 and 0.25–0.37, respectively, in zero-shot prompting. Lower temperature and larger k were associated with higher similarity scores. Conclusions Dynamic, case-matched retrieval improved alignment of LLM-generated CTPA impressions with reference impressions on automated text-similarity metrics. Scores remained moderate, and radiologists’ verification is still required before clinical deployment.
Vera Sorin, Jeremy D. Collins, Lewis Hahn et al.· PLoS ONE· 0 citations
BACKGROUND
Reviewing pathology, imaging, and consultation documents in oncology can be time-consuming, particularly when records originate from external facilities in different file formats. This study aimed to evaluate the impact of a Retrieval-Augmented Generation (RAG)-enabled GPT-4o summarization agent on clinical workflows and quality of outside-record summaries in breast surgical oncology.
METHODS
Initial performance evaluation of a GPT-4o/RAG agent to generate summaries of oncologic reports in 50 charts followed by a prospective pilot test of sequential cases, with each AI summary evaluated using a modified Provider Documentation Summarization Quality Instrument (PDSQI-9; 1-5 Likert scale), including dichotomized ratings (low [1-3], high [4, 5]), binomial testing, frequency and type of user-reported errors, clinician-coded error criticality (treatment-impacting vs noncritical). Pre- and post-use survey of documentation burden (NASA TLX) and user experience was performed.
RESULTS
Among 62 cases, AI-generated summaries were rated high for accuracy, usefulness, succinctness, and source citation. Thoroughness without omission was rated low in 28 (45%) summaries. Errors were noted in 25 (40%) surveys, with 13 (52%) classified as critical (treatment-impacting). The most common error type involved imaging, reported in 17 (68%) cases. For perceived time savings, the median response was neutral, but qualitative feedback described the tool as helpful for straightforward cases and as reducing typing burden but requiring workflow adjustment and improvements for complex cases.
CONCLUSIONS
Although users rated RAG-enabled GPT-4o agent-generated documentation summaries favorably on several quality domains, they frequently lacked thoroughness and occasionally contained treatment-relevant errors. Human review and further iteration of the technology remain necessary before implementation.
Ko Un Park, Bergen K. Sather, A. Shah et al.· Annals of Surgical Oncology· 0 citations
GPT-5 reliably adapts linguistic style to clinical personas but produces limited specialty-specific output diversity, supporting its role as a decision-support adjunct rather than an autonomous specialist simulator.
Derna Stifini, A. Della Penna, André L. Mihaljevic et al.· npj Digital Medicine· 0 citations