Skip to content

CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

External evaluation shows that while citation validity remains strong, evidence utilization, span alignment, and refusal calibration become harder under domain shift, indicating that trustworthy RAG systems require explicit validation between retrieval and final answer delivery.

Abstract

Retrieval-augmented generation (RAG) can improve access to complex information; however, retrieving evidence alone does not ensure that answers are grounded, citation-valid, or appropriately refused. This paper introduces CiteGuard-RAG, a validation-centered AI system for evidence-grounded question answering. The system integrates hybrid semantic-lexical retrieval, citation-constrained generation, sentence-level grounding validation, and single-pass regeneration. Validation is used at runtime to determine whether a candidate answer should be accepted, refused, or regenerated before final delivery. CiteGuard-RAG is evaluated on 400 questions across a controlled housing-law dataset, PrivacyQA, and CUAD. In the controlled evaluation, it achieves 99.1% retrieval accuracy, 98.3% grounded-answer accuracy, and 98.3% citation validity, with no validation-detected hallucinations. Ablation results show that grounded-answer accuracy drops sharply when validation is removed, even when retrieval accuracy remains unchanged. External evaluation shows that while citation validity remains strong, evidence utilization, span alignment, and refusal calibration become harder under domain shift. These findings indicate that trustworthy RAG systems require explicit validation between retrieval and final answer delivery. CiteGuard-RAG provides a practical architecture for linking retrieval, generation, citation checking, abstention, and regeneration in high-stakes information access.

View source

Similar papers

Preprint Aug 2026

Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries

Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing: a system that answers everything outscores one that declines when its knowledge base cannot support an answer. Building on the confidence-target analysis of Kalai et al. (2025), we present a penalty-aware evaluation framework for d...

Alden Do Rosario, Hussein Younes, F. Pires · 1 citation
Open access Aug 2026

Enhanced Hybrid Retrieval-Augmented Model for Question Answering in High-Sensitivity Domains

Arabic question-answering systems in high-sensitivity domains require not only accurate retrieval but also reliable evidence grounding and effective hallucination mitigation, as incorrect or unsupported responses may have serious consequences. Existing retrieval and generation approaches do not fully integrate reliable...

A. Aloqla, Reda Salama, Wajdi Alghamdi et al. · 0 citations
#machine learning Preprint Sep 2026

Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG

Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when instructed to abstain, 12 generators answer 40.0-99.3% of insufficient-evidence questions. Training generators to abstain t...

Su-Ting Chen, Pei-Chun Hua, Yun-Ming Xiao · 0 citations
Sep 2026

Evidence Conflict: Diagnosing and Mitigating Retrieval-Augmented Generation Under Contradictory Evidence

Retrieval-Augmented Generation (RAG) is widely evaluated under the assumption that retrieved passages are jointly compatible, yet retrieval over noisy corpora frequently surfaces passages that disagree in stance, numbers, named entities, or temporal scope. Existing RAG benchmarks measure faithfulness against gold conte...

U. Patole · 0 citations
#natural language process... Preprint Sep 2026

Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification

Analysis reveals that scaling model size, adding reasoning effort, expanding to authoritative web sources, and applying medical fine-tuning do not resolve these failure modes, demonstrating that they represent fundamental limitations of the retrieve-then-verify paradigm in open-ended medical settings rather than artifa...

He-Yuan Huang, Ji Dai, Alexandra DeLucia et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.