It is suggested that RAG LLMs may be supplementary tools for evidence retrieval and synthesis but cannot, at this time, fully replace comprehensive expert review of the medical literature.
Abstract
Background: Large language models (LLMs) that use retrieval-augmented generation (RAG) are increasingly used to answer clinical questions, although the evaluation of these systems remains limited. Building on previous studies conducted by our team, this case report aimed to improve upon this knowledge gap by applying a reusable methodology to compare the performance of eight LLMs that utilize RAG techniques for evidence synthesis. Case Presentation: Eight commercially available RAG LLM tools (OpenEvidence, Undermind, Consensus, SciSpace, Elicit, MediSearch, EvidenceHunt, and Scite) were evaluated using twelve ChatGPT-generated clinical questions on the topics of treatment, etiology, and prognosis. To enable comparison, we prompted ChatGPT to identify all key unique medical concepts from the full set of LLM responses to each question. Concepts were categorized as critical ("must-have") or non-critical ("nice-to-have") for answering the clinical question. Experienced information scientists were consulted at each step for their expertise. Descriptive statistics and Kruskal-Wallis tests were used to compare performance across tools and question categories. No significant differences were found among the eight RAG LLMs in their coverage of "must-have" (p=0.95) or "nice-to-have" (p=0.16) key unique medical concepts, and no single tool consistently captured all identified concepts. Conclusions: These findings suggest that RAG LLMs may be supplementary tools for evidence retrieval and synthesis but cannot, at this time, fully replace comprehensive expert review of the medical literature. The evaluation framework presented here may be a useful model for future comparative assessments of rapidly evolving AI evidence synthesis tools.
Background/Objectives: Large language models (LLMs) show promise for clinical decision support, yet their accuracy in interpreting specialized medical guidelines remains uncertain. Retrieval-augmented generation (RAG) may enhance performance by grounding responses in authoritative knowledge bases. This study aimed to compare the accuracy, comprehensiveness, and safety of RAG-enhanced versus standard LLMs for answering clinical questions derived from the German S3 guideline for oral cavity carcinoma. Methods: We conducted a prospective, single-blind benchmark study evaluating six LLMs: one RAG-enhanced model (Custom GPT with guideline access), one consensus-based model (ConsensusGPT), and four standard models (DeepSeek-V3.2, Mistral Small 3.2, Qwen3-Next-80B, GPT-OSS-120B). Fifty clinical questions covering 17 guideline domains were presented to each model three times, yielding 900 evaluations. Three expert reviewers assessed responses using 5-point Likert scales for accuracy, comprehensiveness, and clarity, under a single-blind procedure, the effectiveness of which was tested by a pre-specified manipulation check. We then ran a paired within-model experiment in which each base model was queried with and without guideline access through a transparent, openly released retrieval pipeline, and scored every response with a condition-blind automated judge alongside deterministic retrieval metrics computed from the logs. Secondary outcomes included hallucination rates and guideline citation behavior. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs). Results: In a paired within-model design that held each base model fixed, adding transparent guideline retrieval improved accuracy—significantly in the three weaker open-weight models (Mistral, Qwen3, and GPT-OSS) and directionally in the already-strong DeepSeek and GPT-5 bases. Because a pre-specified blinding check found that experts could still identify retrieval-augmented answers with 98.5% accuracy, we anchored causal interpretation on measures that do not depend on the human raters, ranked by their independence: deterministic, log-derived retrieval metrics first, and then an automated, condition-blind LLM judge, whose agreement with the experts (Spearman ρ = 0.81, 95.7% within-one agreement) establishes shared calibration rather than independence from their bias. Deterministically from the retrieval logs, citation groundedness rose from 0% to 51–89% and retrieval recall@5 was 92%. On the judge, content-level hallucination fell from 42% to 4% and accuracy rose by a pooled +0.64 points (95% CI 0.47–0.80); the accuracy gain persisted after adjustment for response length (+0.48, 95% CI 0.22–0.73), which retrieval shortened rather than lengthened. The accuracy gain was large for weaker base models and small or non-significant for already-strong ones, whereas the hallucination and auditability gains were consistent across all models. The human ratings reproduced the judge’s accuracy effect (+0.61, 95% CI 0.49–0.74), and GPT-5 run through the transparent pipeline showed no significant difference from the proprietary Custom GPT (judge accuracy 4.48 vs. 4.58). Conclusions: Guideline retrieval yields a reproducible, largely base-independent improvement in the safety and auditability of LLM answers to clinical guideline questions, with accuracy gains concentrated in weaker base models. Because retrieval-augmented answers are recognizable to experts, rigorous evaluation should rely on rater-independent measures, and residual hallucination continues to require human oversight.
Andreas Vollmer, Lara Schorn, Felix Schrader et al.· Diagnostics· 0 citations
Background/Objectives: Clinical documentation places a significant time burden on healthcare professionals, including in the context of home care. Large language models (LLMs) offer potential for automated note generation, but current approaches rely on static prompt templates that fail to generalize across care settings, languages, and documentation formats. This study proposes and evaluates an adaptive retrieval-augmented generation (RAG) framework that uses retrieval as a format adaptation mechanism, enabling the generation of structured clinical notes from patient–provider transcripts across various documentation formats without model fine-tuning. Methods: The proposed framework retrieves dialogue–note pairs that demonstrate the structure of specific sections, allowing the transfer of formatting knowledge during inference. Experiments were conducted on three datasets covering two languages and various documentation formats: the Japanese Visiting Nurse corpus (JP-VN), MTS-Dialog, and ACI-BENCH. Six controlled conditions were evaluated: zero-shot prompting (C1), static few-shot prompting (C2), dense retrieval (C3), random retrieval (C4), sparse BM25 retrieval (C5), and hybrid retrieval using reciprocal rank fusion (RRF) (C6). Performance metrics include structural adherence to required section headings, content quality (ROUGE-1, BLEU, BERTScore), and the number of hallucinated clinical entities per generated record. Results: Structure compliance increased from 0–37% under static conditions (C1/C2) to 91–100% under all adaptive RAG conditions (C3–C6) across all datasets. On MTS-Dialog, dense retrieval achieved the highest content quality (ROUGE-1: 0.519 vs. 0.446–0.492 for C4–C6; p<0.001). Hallucinated entities in JP-VN decreased from 2.73–3.58 per note (C1/C2) to 1.15–1.30 (C3–C6), an approximately 55–56% reduction. Conclusions: Adaptive RAG can improve structure compliance and reduce hallucinations in multilingual clinical note generation without dataset-specific prompt engineering or model fine-tuning. These findings support retrieval-based format adaptation as a generalizable mechanism for diverse clinical documentation contexts.
Whether contemporary LLMs can reproduce the research outcomes of a fully documented human study: a 1991 article that identified dermatophytosis (ringworm) in historical fine art was evaluated.
The proposed framework offers a practical and scalable approach to mitigating hallucinations without requiring task-specific fine-tuning, highlighting the potential of retrieval-augmented approaches for trustworthy artificial intelligence (AI)-assisted healthcare applications.
PURPOSE
To evaluate whether large language models (LLMs) can provide accurate, complete, and audience-adapted answers to common spine-surgery-related questions for patients and family practitioners.
METHODS
Ten frequently asked spine-surgery questions were collected at a level 1 trauma center and simplified linguistically. Five LLMs (ChatGPT, Claude 3.5 Sonnet, Gemini Advanced 1.5 Pro, Copilot Pro, and DeepSeek V3) were queried using zero-shot prompting with persona-specific instructions for family practitioners and middle-aged patients. Responses were assessed by spine surgeons and non-medical raters for correctness, completeness, adaptability, and empathy using five-point Likert scales. Readability was quantified using the Flesch Reading Ease Score (FRES).
RESULTS
All LLMs generated largely correct and usable responses. ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers. Gemini and Copilot achieved superior readability and empathy for patient-facing responses. DeepSeek demonstrated balanced performance across all domains. Readability differed substantially between practitioner- and patient-oriented outputs.
CONCLUSION
LLMs can support communication and education following spine surgery when used with structured prompting. Clinical oversight remains essential to mitigate risks related to inaccuracies and hallucinations.
LEVEL OF EVIDENCE
III.
S. Wegmann, T. Rosenkranz, Philipp Egenolf et al.· European spine journal· 0 citations