Skip to content
Review Open access

Narrative Consistency in Large Language Model-Generated Stories: A Survey

2026 · IEEE Access · Vol 14, pp. 105794-105819 · 0 citations · 124 references
Computer Science

TL;DR

This survey examines the problem as narrative consistency, defined as the task-conditioned preservation of binding propositions in the operative narrative state, and introduces a four-category, fourteen-subtype taxonomy comprising World and Setting, Character-Agentive, Event-Structural, and Narration and Discourse categories.

Abstract

Large language models have advanced long-form story generation, yet the resulting narratives often fail to preserve facts, states, and relationships established earlier in the same text. This survey examines the problem as narrative consistency, defined as the task-conditioned preservation of binding propositions in the operative narrative state. Using a targeted evidence-mapping strategy, we analyze 90 papers, with a literature search cutoff of April 23, 2026, covering long-form story generation, narrative evaluation, consistency detection, and mitigation strategies. We distinguish narrative consistency from factuality, faithfulness, surface coherence, and hallucination by centering the dynamically accumulated evidence of the story itself. The survey organizes consistency judgments around three evidence sources, namely story-internal propositions, source-or-canon evidence, and external-world knowledge, as determined by task conditions that specify which source is binding. We introduce a four-category, fourteen-subtype taxonomy comprising World and Setting, Character-Agentive, Event-Structural, and Narration and Discourse categories. We use the taxonomy as a common reference frame for analyzing benchmark coverage, detection methods, mitigation strategies, and open challenges, highlighting where current work concentrates and where coverage remains thin. We also separate legitimate creative extension from task-inconsistent additions that contradict, revise, or exceed the permitted generation setting. The survey closes by identifying needs for evidence-grounded oracles, subtype-aware calibration, omission and discourse-level failure evaluation, and task-conditioned verification.

Read PDF

Similar papers

Preprint Jul 2026

Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction

Long-form fiction writers need memory that answers multi-hop questions about evolving story state: who knows a secret and when they learned it, whether an event preceded the narration that revealed it, whether a setup paid off, and how a relationship shifted. General-purpose retrieval and agent-memory systems represent entities and facts but not the narratological structure these questions turn on, so they surface the wrong evidence or none at all. We introduce the Narrative World Model (NWM), a writer-memory system that pairs a narratology-grounded typed temporal-state graph with query-conditioned hybrid retrieval. To measure memory rather than the answerer, we read every system through a single held-constant Opus 4.8 reader over only that system's chapter-safe evidence, on a reproducible public corpus and a validated multi-hop benchmark, and we compare against the strongest existing temporal-knowledge-graph agent-memory framework, Graphiti/Zep (Rasmussen et al., 2025). NWM substantially and significantly outperforms this baseline on multi-hop narratological QA across both corpora, and far exceeds GraphRAG and flat retrieval. The advantage is representational rather than an artifact of extraction: it survives rebuilding the baseline with NWM's own extractor, and traces to its narratology-grounded structure and query-conditioned retrieval, not to graph size or extractor quality.

M. Saifullah, Thomas Kornmaier, Taaha Kazi et al. · 1 citation · ⚡1
Open access 2026

COGNAC at SemEval-2026 Task 4: Evaluating Narrative Components with LLMs for Hard Story Similarity Cases

This system for the Narrative Similarity task at SemEval-2026 (Task 4), where the goal is to determine which of two candidate stories is more similar to an anchor story directly or via vector representations, finds that chain-of-thought–style prompting with detailed reasoning outputs achieves comparable results to the scoring approach on difficult examples.

Tisa Islam Erana, Azwad Anjum Islam, Anshu Kiran Sharma et al. · 1 citation · ⚡1
Preprint Aug 2026

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video, is introduced, a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal.

Yuheng Huang, Jianlang Chen, Jiayang Song et al. · 0 citations
Conference Open access Jul 2026

Structure-Grounded Document Agents for Faithful Long-Context Reasoning

Large language models (LLMs) struggle with longdocument reasoning: naively packing entire documents into the context window leads to degraded performance, while retrievalaugmented generation (RAG) fragments documents into isolated chunks that lose structural coherence. Moreover, even when existing agents produce correct answers, their intermediate reasoning steps are often unfaithful to the retrieved evidence. We propose StructAgent, a document-oriented LLM agent that integrates structure-aware navigation, sequential reading, and evidence-constrained generation into a unified agentic loop. StructAgent first parses the document hierarchy (sections, tables, and cross-references), then navigates this structure tree to locate relevant regions, reads contiguous passages to preserve local context, and finally generates answers whose reasoning chains are explicitly grounded in cited evidence. We evaluate StructAgent on three long-document benchmarks-QASPER, HotpotQA, and QuALITY-measuring both answer accuracy (F1) and reasoning faithfulness via two newly introduced metrics: Evidence-Reasoning Consistency (ERC) and Reasoning-Answer Consistency (RAC). Experimental results show that StructAgent achieves absolute F1 gains of 3.1-10.9 points depending on the baseline and dataset, along with higher faithfulness scores compared to vanilla RAG, long-context LLMs, and ReAct-based agents, especially on the length-controlled QASPER setting.

Jingfeng Zhou, Zhizhen Zhu, A. Zhu et al. · 1 citation
Preprint Aug 2026

CraftAlign: Feature-Grounded Evaluation and Revision Guidance for AI Stories

Large language models can now generate fluent and complete stories, yet many outputs still feel formulaic and unnatural because of cliches, over-explanation, linear causal progression, and stereotyped endings, an immediately recognizable AI flavor. Existing detection and evaluation methods often stop at source labels or holistic scores, while revision methods typically target predefined issues through localized edits, limiting their ability to support multiple plausible revision strategies or guide story-wide changes in information release, causal organization, and ending treatment. We introduce CraftAlign, a framework that aligns AI stories with the craft of human storytelling by both assessing Human/AI writing patterns and providing revision guidance. CraftAlign comprises two learned modules and an inference-time guidance pipeline. A feature estimator built on Qwen3.5-9B predicts 304 explicit writing features spanning style and narrative. A class-conditional energy model scores the resulting feature configuration against Human and AI writing patterns, conditioning on the original writing prompt when available. At inference time, CraftAlign applies schema-valid structured perturbations, selects changes that move the feature configuration toward the Human writing pattern, and converts them into natural-language guidance for a separate editor to rewrite the full story. Experiments show that CraftAlign accurately distinguishes Human and AI writing patterns and that its guidance outperforms revision baselines across editors and in a human study.

Yang Yang, Boyun Xu, Shaofeng Liang et al. · 0 citations
Preprint Jul 2026

The Story Shapes the Agent: Narrative Priors in LLM Behavior

Persona prompting is widely used to steer LLM agent behavior, yet the narrative framing of a task can matter more than the assigned persona. We isolate this effect through structural isomorphism, constructing three text-based investigation games that share the same action space, stage progression, and resource constraints while varying only task narrative: disease investigation, IT troubleshooting, and murder mystery. Across 1,890 sessions spanning 3 models and 10 personas, we identify narrative priors: systematic action tendencies activated by a task's story framing, independent of its decision structure. Narrative priors explain 5-31x more behavioral variance than persona, are consistent across model architectures, and in two of three domains are negatively associated with task success. Persona effects that do transfer across narratives arise from behavioral anchors, persona descriptions whose language maps directly onto shared actions. Causal interventions confirm this: removing anchor words from a high-transfer persona reduces cross-narrative consistency by 95%. Our framework also generalizes to a held-out fourth narrative and yields a persona-selection method that improves cross-narrative transfer. These results suggest that LLM behavior that survives narrative changes should be grounded in concrete actions rather than abstract descriptions.

Yixuan Wang, James C. Lester, Shashank Srivastava · 0 citations