Comparing human narrative engagement with model attention mechanisms suggests explanations for degraded narrative comprehension and targets for future development.
Abstract
Although LLM context lengths have grown, there is evidence that their ability to integrate information across long-form texts has not kept pace. We evaluate one such understanding task: generating summaries of novels. When human authors of summaries compress a story, they reveal what they consider narratively important. Therefore, by comparing human and LLM-authored summaries, we can assess whether models mirror human patterns of conceptual engagement with texts. To measure conceptual engagement, we align sentences from 150 human-written novel summaries with the specific chapters they reference. We demonstrate the difficulty of this alignment task, which indicates the complexity of summarization as a task. We then generate and align additional summaries by nine state-of-the-art LLMs for each of the 150 reference texts. Comparing the human and model-authored summaries, we find both stylistic differences between the texts and differences in how humans and LLMs distribute their focus throughout a narrative, with models emphasizing the ends of texts. Comparing human narrative engagement with model attention mechanisms suggests explanations for degraded narrative comprehension and targets for future development. We release our dataset to support future research.
NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video, is introduced, a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal.
Yuheng Huang, Jianlang Chen, Jiayang Song et al.· 0 citations
This survey examines the problem as narrative consistency, defined as the task-conditioned preservation of binding propositions in the operative narrative state, and introduces a four-category, fourteen-subtype taxonomy comprising World and Setting, Character-Agentive, Event-Structural, and Narration and Discourse categories.
Keunhyeung Park, Seunguk Yu, Jinhee Jang et al.· IEEE Access· 0 citations
The study identifies prompt anchoring as a source of methodological variation in LLM-assisted content analysis, indicating that anchoring strategies should be explicitly specified, justified, and reported as part of the study methodology.
Initial human feedback reveals that AI-written novels contain interesting descriptions and concepts, but often fail in long-range coherence and prose quality, including conceptual repetition, distracting details, and weak dialogues.
It is suggested that long-form narratives can emerge from explicit world-state tracking and goal-driven multi-agent generation, providing a foundation for controllable and structurally coherent long-form narrative generation.
Aayush Aluru, C. Ho, Muhammad Hammouri et al.· 0 citations
This system for the Narrative Similarity task at SemEval-2026 (Task 4), where the goal is to determine which of two candidate stories is more similar to an anchor story directly or via vector representations, finds that chain-of-thought–style prompting with detailed reasoning outputs achieves comparable results to the scoring approach on difficult examples.
Tisa Islam Erana, Azwad Anjum Islam, Anshu Kiran Sharma et al.· SemEval@ACL· 1 citation· ⚡1
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.