Evidence Conflict: Diagnosing and Mitigating Retrieval-Augmented Generation Under Contradictory Evidence
Abstract
Retrieval-Augmented Generation (RAG) is widely evaluated under the assumption that retrieved passages are jointly compatible, yet retrieval over noisy corpora frequently surfaces passages that disagree in stance, numbers, named entities, or temporal scope. Existing RAG benchmarks measure faithfulness against gold contexts, and fact-verification corpora evaluate classifiers rather than generators; neither captures how a generator should behave when its evidence is internally contradictory. To address these challenges, we propose EvidenceConflict, a controlled benchmark spanning twelve evidence profiles built from VitaminC and Climate-FEVER, together with a four-axis metric suite that separates conflict detection, acknowledgment, citation balance, and claim-level faithfulness. We further propose ConflictDetector, a lightweight pre-generation predictor that combines pairwise NLI with entity, number, and date mismatch heuristics through a calibrated logistic regression, gating a conflict-aware prompt template. Experiments on a production-grade closed-source LLM and a deterministic template baseline show that standard prompting collapses to a single stance under contradictory evidence, while conflict-aware variants substantially reduce overconfidence, raise acknowledgment, and improve citation balance without harming faithfulness on non-conflicting cases. Ablations further confirm the contribution of each detector signal.