Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research
An evaluation method is proposed that pinpoints which agent introduced each error by locally testing agent invocations for faithfulness and verifiability relative to their own inputs and proposes a four-type taxonomy to categorize the discovered errors: hallucination, uncited input reliance, uncited output, or insufficient citations.