Skip to content
Preprint

When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

Aug 2026 · 1 citation · 31 references
Computer Science

TL;DR

WSE-bench is introduced, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling, showing that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities.

Abstract

Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branching, player-shaped trajectories develop. Across frontier models, Consistency and Richness do not form a smooth trade-off: their empirical Pareto frontier is non-concave, with several non-dominated intermediate configurations that no positive linear weighting can select. Added structure can enrich trajectories, but it does not uniformly improve coherence and may shorten them. Model scale chiefly improves sustained generation, without producing reliable gains in canonical coherence or meaningful development. These results show that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities. WSE-bench makes those dynamics visible by extending narrative evaluation from finished stories to the processes that create them.

View source

Similar papers

Preprint Aug 2026

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

This work forms this challenge as Narrative Commitment Preservation (NCP), and introduces NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses that each environment includes a structured narrative specification that can automatically check throughout the interaction between the player agent...

Yingpeng Ma, Jianhao Yan, Bei-Ning Shi et al. · 1 citation
#natural language process... Preprint Sep 2026

MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing

Vibe Narrativizing is formulated as turning natural-language writing requirements into a finished story, and MUSE, a Theory-Harnessed Story Engine, addresses two bottlenecks: rule quality and sustained rule realization.

Jian-Xiang Ma, Xiaocui Yang, Da-Ling Wang et al. · 0 citations
Preprint Sep 2026

HappyWorld-Bench

HappyWorld-Bench is introduced, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them, and highlights the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.

Zhi-Qi Bai, Ju-Nai Cai, Yi-Xin Chen et al. · 0 citations
Preprint Sep 2026

Code Plans, Diffusion Renders: Open-Ended Generative World Modeling

Extensive experiments demonstrate that the CoDeR framework substantially extends the capabilities of existing world models, enabling long-term memory, open-ended interactions, autonomous evolution, and persistent multi-agent dynamics, while achieving state-of-the-art performance across multiple evaluation settings.

Zi-Xun Fang, Ya-Wen Shao, Kai Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

CivBench is presented, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP), and two interface-level metrics are introduced that the environment makes measurable: Proactive Monitoring Rate (PMR) and RAG@10, capturing whether c...

Austin Andrews, L. Wilkinson, Jamie Heagerty et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.