Skip to content
Preprint

When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim Verification

Aug 2026 · 0 citations · 50 references
Computer Science

TL;DR

Bio-GRACE shows that retrieval utility is source-dependent, motivates selective retrieval, and exposes why retrieval recall and lexical evidence overlap are insufficient for biomedical fact-checking, and introduces Bio-GRACE, a gold-reference-normalized diagnostic for measuring whether retrieved evidence recovers the decision benefit of reference evidence.

Abstract

Biomedical fact-checking systems must do more than predict whether a claim is supported, contradicted, or unaddressed: they should also produce evidence that is faithful, complete, and useful for verification. We study this evidence-generation setting on CARE-XAI, a unified benchmark spanning five biomedical and health fact-checking sources. We compare base instruction LLMs, PubMed retrieval-augmented LLMs, fine-tuned LLMs, label-only LLMs, and biomedical encoder classifiers under a shared evaluation protocol. Biomedical classifiers remain strongest for verdict-only prediction, while fine-tuned LLMs are the strongest evidence-generating systems. PubMed retrieval is mixed: it helps PubMed-aligned sources such as PubMedQA and SciFact, but can distract models on broader public-health claims. We introduce Bio-GRACE, a gold-reference-normalized diagnostic for measuring whether retrieved evidence recovers the decision benefit of reference evidence. Bio-GRACE shows that retrieval utility is source-dependent, motivates selective retrieval, and exposes why retrieval recall and lexical evidence overlap are insufficient for biomedical fact-checking.

View source

Similar papers

Preprint Aug 2026

Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search

The Evidence-Grounded Group Relative Policy Optimization (EG-GRPO) is proposed to perform reinforcement learning on BioCheck Agent with a task-specific reward that incentivizes advanced search behavior and high-quality evidence retrieval while penalizing hallucinations.

Jiongxiao Wang, Di Ma, Chaoqun Ni · 0 citations

EviGen: Predictive Evidence Scaffolding for Verifiable Clinical Rationale Generation

Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed to reason about likely clinical outcomes. However, comprehensive clinician review of these records is impractical, and LLM-based processing is costly and often unreliable...

Feng-Nan Li, Heman Burre, Li-Wen Sun et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine

Medical knowledge evolves continuously, whereas the parametric knowledge encoded in large language models (LLMs) is fixed at training time. External retrieval, including retrieval-augmented generation (RAG), can provide access to newly available evidence, but retrieved information may be irrelevant, incomplete, or conf...

Shuai Wang, Yi-Ze Zhao, Qing-Yu Chen · 0 citations
#natural language process... Preprint Aug 2026

Quantitative Evidence Mining for Plausibility-Aware Biomedical AI

This work outlines a framework for plausibility-aware AI that treats extracted claims not as final answers but as auditable evidence objects, making clear what was measured, how much it changed, in which setting, with what uncertainty, and from which source.

N. S. Babaiha, Stefan Geißler, Marie-Christine Simon et al. · 0 citations
Conference Aug 2026

MedEvidence-RAG for Joint Evidence Retrieval and Generator Alignment in Medical Visual Question Answering and Report Generation

Medical vision-language models remain vulnerable to unsupported findings because their outputs are weakly grounded in images and external clinical evidence. We propose MedEvidence-RAG, which couples a mixture-of-experts (MoE) multimodal retriever with retrieval-conditioned generator alignment. The retriever uses a SigL...

Xing-Yu Qu, Shuang-Quan Li, Liu Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges

We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality human judgments are scarce. First, we augment existing human-labelled biomedical benchmarks with deterministic, metric-grounded mutations that produce auditable preference pairs. Second, we evaluate judges beyond ag...

R. D. de Oliveira, Federico Pittino, J. Gwinnutt et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.