Skip to content
Preprint

Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research

Aug 2026 · 0 citations · 24 references
Computer Science

TL;DR

An evaluation method is proposed that pinpoints which agent introduced each error by locally testing agent invocations for faithfulness and verifiability relative to their own inputs and proposes a four-type taxonomy to categorize the discovered errors: hallucination, uncited input reliance, uncited output, or insufficient citations.

Abstract

Deep research (DR) systems produce long-form cited reports by orchestrating multiple agents that search and synthesize information from the web. Citations are the primary mechanism for evaluating the faithfulness of these reports, yet current DR systems exhibit poor citation recall. Moreover, improving citation recall is challenging because DR systems are complex multi-agent architectures where information passes through agents like a telephone game, and both content and citations can get corrupted along the way. We propose an evaluation method that pinpoints which agent introduced each error by locally testing agent invocations for faithfulness and verifiability relative to their own inputs. Furthermore, we propose a four-type taxonomy to categorize the discovered errors: hallucination, uncited input reliance, uncited output, or insufficient citations. Applying our method to three top-ranked open-source DR systems, we obtain actionable diagnostics. Almost every agent makes a lot of mistakes with the exception being those that summarize a single document. We find that the dominant error type varies systematically across agents, where the orchestrator mistakes are mostly citation-related. We find that 84.7% of final-report errors in AI-Q originate at the orchestrator, roughly 31% of them hallucinations and the rest citation mistakes. Guided by these insights, we demonstrate that two simple interventions raise citation recall by 5% without degrading output quality.

View source

Similar papers

#natural language process... Preprint Sep 2026

ReCite: Agentic Reasoning for Faithful Citation

Accurate citations are the foundation of academic writing, tracing intellectual origins and substantiating core claims. However, manually navigating the growing volume of scientific literature is increasingly difficult, prompting reliance on automatic citation recommendation. While modern retrieval-augmented architectures have largely mitigated the fabrication of non-existent papers, current systems relying on semantic similarity struggle with misattribution, often citing authentic papers that fail to logically support the author's claim. To address this challenge, we argue that accurate citation requires a shift from similarity-based search to active, claim-level reasoning. We propose ReCite, a decoupled agentic framework that orchestrates location perception, intent-aware query planning, and reflective verification. Trained on synthesized reasoning trajectories, our agent verifies claim-evidence consistency and triggers self-correction loops when retrieved candidates lack logical support. Experiments demonstrate that our lightweight framework outperforms state-of-the-art massive generative models in strict citation accuracy. By grounding literature matching in verifiable logic rather than semantic overlap, ReCite establishes a reliable foundation for automated academic writing.

Yu-Yang Huang, Bobo Li, Jia-Jia Song et al. · 0 citations
Review Aug 2026

Who Are You Explaining To? A Multi-Agent System for Audience-Aware XAI Narratives

This work introduces XstrAI, an audience-aware multi-agent framework that treats local explanations as fixed evidence and structures how it is communicated to each target reader, and evaluates XstrAI on diabetes and stroke risk prediction against 11 baselines.

F. Musicco, Danilo Danese, Giuseppe Fasano et al. · 0 citations
#natural language process... Preprint Aug 2026

Lazy Grounding: Attacking Search Agents with Factual Evidence

This work exposes lazy grounding by injecting nearby evidence from answer-changing rewrites of benchmark questions into the search corpora, and shows that robust search agents must defend against not only misinformation but also the misapplication of factual evidence.

Yu-Lin Zhang, Yukun Huang, Sanxing Chen et al. · 0 citations
Preprint Aug 2026

ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives

Experiments across multiple long-context narrative question answering and claim verification settings show that ClueWeaver substantially improves local end-to-end language models while providing evidence coverage and paragraph-referenced reasoning traces.

Ji-Hao Zhu, Zhi-Wei Yang, Wen-Xiao Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Redesigning and Auditing Deep Research Writing for Faithful Reports

CLAIMPROBE is introduced, a claim-level audit that decomposes DR reports into claims and measures hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence and proposes CLAIMWRITER, a hierarchical claim-based writer that extracts source facts, maps them to a query-derived outline, and drafts each section from a source-linked claim representation.

Hiroaki Hayashi, P. Venkit, Prafulla Kumar Choubey et al. · 0 citations
#natural language process... Preprint Aug 2026

You Know What I Mean: A Benchmark for Agentic Conversational Reference Grounding

The results show that CoRG remains challenging for current agents, even the best agent reaches only 67.0% success rate, leaving one third of references unresolved, and position CoRG as a concrete benchmark for studying how agents search, inspect, and verify information in realistic multi-tool environments.

Karen Fuchs, Uri Katz, Yoav Goldberg · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.