Skip to content
Preprint

Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search

Aug 2026 · 0 citations · 47 references
Computer Science

TL;DR

The Evidence-Grounded Group Relative Policy Optimization (EG-GRPO) is proposed to perform reinforcement learning on BioCheck Agent with a task-specific reward that incentivizes advanced search behavior and high-quality evidence retrieval while penalizing hallucinations.

Abstract

Automated fact-checking is essential for ensuring the reliability of public health information, yet the biomedical domain poses unique challenges. Validating biomedical claims requires rigorous interpretation of scientific literature, assessment of retrieved evidence, and comprehensive justification toward the conclusion. Although Large Language Models (LLMs) enhanced by Retrieval-Augmented Generation (RAG) and agentic search perform automated fact-checking in a retrieve-then-verify paradigm, current methods still output isolated prediction labels, lacking explanatory depth and offers limited utility for human understanding. To bridge this gap, we introduce an LLM-based agent named BioCheck Agent that generates structured biomedical fact-checking reports with agentic search. Rather than merely outputting supported or refuted labels, our agent synthesizes final conclusions with retrieved evidence and rigorous analysis. To ensure domain-specific accuracy, BioCheck Agent exclusively searches high-quality scientific literature in PubMed, utilizing advanced Boolean search operators. Recognizing that direct prompting often results in hallucinations and low-quality reports, especially for lightweight open-source models, we further propose the Evidence-Grounded Group Relative Policy Optimization (EG-GRPO) to perform reinforcement learning on BioCheck Agent with a task-specific reward that incentivizes advanced search behavior and high-quality evidence retrieval while penalizing hallucinations. Our experimental results show that compared to the base model Qwen3.5-4B, BioCheck Agent with EG-GRPO improves label prediction accuracy on SciFact by 9.95%. Furthermore, it achieves a 3.7% higher evidence quality score and a 19.63% lower evidence hallucination rate, demonstrating its ability to generate biomedical fact-checking reports with improved accuracy and quality.

View source

Similar papers

Preprint Aug 2026

When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim Verification

Bio-GRACE shows that retrieval utility is source-dependent, motivates selective retrieval, and exposes why retrieval recall and lexical evidence overlap are insufficient for biomedical fact-checking, and introduces Bio-GRACE, a gold-reference-normalized diagnostic for measuring whether retrieved evidence recovers the d...

Pritam Deka, Prabhjot Singh · 0 citations

EviGen: Predictive Evidence Scaffolding for Verifiable Clinical Rationale Generation

Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed to reason about likely clinical outcomes. However, comprehensive clinician review of these records is impractical, and LLM-based processing is costly and often unreliable...

Feng-Nan Li, Heman Burre, Li-Wen Sun et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine

Medical knowledge evolves continuously, whereas the parametric knowledge encoded in large language models (LLMs) is fixed at training time. External retrieval, including retrieval-augmented generation (RAG), can provide access to newly available evidence, but retrieved information may be irrelevant, incomplete, or conf...

Shuai Wang, Yi-Ze Zhao, Qing-Yu Chen · 0 citations
Conference Aug 2026

ReCLLaMA: A Reasoning-Centered LLM Agent for Medical Diagnosis

Large Language Models (LLMs) have demonstrated impressive capabilities in natural language understanding, yet their application to clinical diagnosis remains constrained by hallucinations, limited interpretability, and the absence of explicit reasoning mechanisms. Recent AI agents extend LLMs beyond passive text genera...

Yang Zhao, Sai-Yun Dong, Xing-Hua Shi · 1 citation
#natural language process... Preprint Aug 2026

Quantitative Evidence Mining for Plausibility-Aware Biomedical AI

This work outlines a framework for plausibility-aware AI that treats extracted claims not as final answers but as auditable evidence objects, making clear what was measured, how much it changed, in which setting, with what uncertainty, and from which source.

N. S. Babaiha, Stefan Geißler, Marie-Christine Simon et al. · 0 citations
Preprint Aug 2026

An Evidence-Grounded Retrieval-Augmented Transformer Framework for Health Misinformation Verification

The proposed framework provides a practical foundation for developing context-aware and evidence-driven health misinformation verification systems for Nigeria and other resource-constrained settings and highlights the importance of comprehensive and authoritative knowledge sources for reliable health misinformation ver...

Isah M. Bukar, Bala Mairiga Abduljalil, Bashir Saleh Maina et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.