Skip to content

CAGE: Cognitive Attribution Graphs for Faithful Inline Citation Generation in Long-Form Question Answering

Jul 2026 · arXiv.org · Vol abs/2607.24236 · 1 citation · 37 references
Computer Science

TL;DR

This work proposes CAGE (Cognitive Attribution Graphs for Citation Generation), a two-stage framework that introduces an explicit cognitive attribution map before answer generation, and demonstrates the effectiveness of attribution-space contraction and map-guided citation generation.

Abstract

Long-form question answering increasingly relies on retrieved evidence to make LLM outputs verifiable, with inline citations tracing claims to source documents. However, existing systems often attach citations that are topically related but insufficient to support their claims. We identify attribution ambiguity as a structural challenge: end-to-end generation must implicitly resolve combinatorial claim--document assignments, obscuring evidential boundaries and increasing the risk of evidence-boundary overrun, where claims exceed cited support. To address this challenge, we propose CAGE (Cognitive Attribution Graphs for Citation Generation), a two-stage framework that introduces an explicit cognitive attribution map before answer generation. CAGE first trains a plug-and-play Cognitive Map Induction Model to construct answer-centered support subgraphs, aligning each semantic answer unit with supporting documents through explicit relations. A Structured Citation Reasoning Model then realizes these units as sentence-level claims with map-aligned citations. Experiments on ASQA, ELI5, and ExpertQA show that CAGE achieves state-of-the-art performance, demonstrating the effectiveness of attribution-space contraction and map-guided citation generation.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

AtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page Documents

Large language models answering questions over multi-page documents are expected to cite the supporting pages, yet supplied citations are sometimes inaccurate, and current evaluations score citations at generation time or against text passages: no existing benchmark evaluates whether a system can verify and correct a p...

Chen Qian, Yi-Meng Wang, Yu Chen et al. · 0 citations
#natural language process... Preprint Sep 2026

ReCite: Agentic Reasoning for Faithful Citation

Accurate citations are the foundation of academic writing, tracing intellectual origins and substantiating core claims. However, manually navigating the growing volume of scientific literature is increasingly difficult, prompting reliance on automatic citation recommendation. While modern retrieval-augmented architectu...

Yu-Yang Huang, Bobo Li, Jia-Jia Song et al. · 0 citations
Preprint Aug 2026

Data Citation for Large Language Models: A Challenge

It is argued that data citation for large language models is an open challenge, distinct from document-level citation grounding and harder to solve.

G. Silvello · 0 citations
Preprint Aug 2026

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering

Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by-products in which required organizational relations remain implicit across heterogeneous sources. Existing benchmarks provide realistic multi-source evidence, but of...

Akrin Zheng, Alexander Wu, Alaia Liu · 1 citation · ⚡1
#artificial intelligence Preprint Aug 2026

Embedding Models for Stance-Aware Argument Retrieval

This paper introduces diagnostic word-ablation metrics to quantify this phenomenon and proposes a data-centric solution that can alleviate the observed overcorrection in stance-aware argument retrieval and demonstrates that, for sufficiently powerful models, this approach can alleviate the observed overcorrection.

Angelo Sparacino, Francesca Toni, Adam Dejl · 0 citations
Review Open access 2026

Large Language Models for Citation Context and Cited Content Recognition: From Boundary Detection to Evidence Grounding

Citation analysis has traditionally been organized around citation intent, function, stance, citation context analysis, and cited text span identification. Recent large language models (LLMs) have been applied to these tasks through prompting, in-context learning, fine-tuning, annotation assistance, and retrieval-augme...

Shu-Qiao Yang, Xiao-Lei Ma, Ping He · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.