Skip to content
Preprint

Data Citation for Large Language Models: A Challenge

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

It is argued that data citation for large language models is an open challenge, distinct from document-level citation grounding and harder to solve.

Abstract

Large language models increasingly mediate access to information, and a growing body of work asks whether they cite the sources behind their outputs. That work treats citation as a verification device and applies it to textual documents. Scholarly citation serves two further functions, credit and provenance, and it applies to data as much as to text. This paper argues that data citation for large language models is an open challenge, distinct from document-level citation grounding and harder to solve. We ask how such models should cite data so that outputs stay verifiable, provenance stays traceable, and credit reaches data creators and curators. We set out three research directions. Training data attribution has to turn influence estimates into references for corpora absorbed into model parameters. Data citation at inference time has to identify datasets, subsets, and query results at the right granularity and fixity. Citing knowledge graph facts has to define what a reference to a single triple denotes and how credit propagates along provenance. Progress on all three depends on joint work across the database, information retrieval, knowledge representation, and artificial intelligence communities.

View source

Similar papers

#small language model Conference Open access Sep 2026

Improving legal reasoning reliability of large language models through citation resolution enhancements

This project seeks to create the infrastructure needed to make legal citation in OpenJustice rigorous and transparent through various enhancements to the extraction and connection of case metadata, and fully integrating the resulting citation graph into the reranker, maximizing LLM reasoning capabilities.

Aaron Elliott · 0 citations
Preprint Sep 2026

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

Scientific papers require models to integrate evidence across text, equations, figures, tables, code, and datasets while preserving its provenance. Beyond answer correctness, scientific reading requires verifiable outputs from operations such as evidence localization, definition extraction, and consistency checking. We...

Shenxi Wu, Yu-Hong Liu, Hao-Song Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page Documents

Large language models answering questions over multi-page documents are expected to cite the supporting pages, yet supplied citations are sometimes inaccurate, and current evaluations score citations at generation time or against text passages: no existing benchmark evaluates whether a system can verify and correct a p...

Chen Qian, Yi-Meng Wang, Yu Chen et al. · 0 citations
Review Open access 2026

Large Language Models for Citation Context and Cited Content Recognition: From Boundary Detection to Evidence Grounding

Citation analysis has traditionally been organized around citation intent, function, stance, citation context analysis, and cited text span identification. Recent large language models (LLMs) have been applied to these tasks through prompting, in-context learning, fine-tuning, annotation assistance, and retrieval-augme...

Shu-Qiao Yang, Xiao-Lei Ma, Ping He · 0 citations
Open access Sep 2026

LLM-assisted writing and citation advantage: evidence from scientific publications before and after ChatGPT release

Generative artificial intelligence has become a routine part of academic writing. While much of the debate has focused on questions of integrity and authorship, less attention has been paid to how AI-assisted writing may affect research evaluation itself. This paper asks a straightforward but important question: does t...

S. Paklina, P. Parshakov, Elena Rapoport · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.