WSE-bench is introduced, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling, showing that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities.
Abstract
Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branching, player-shaped trajectories develop. Across frontier models, Consistency and Richness do not form a smooth trade-off: their empirical Pareto frontier is non-concave, with several non-dominated intermediate configurations that no positive linear weighting can select. Added structure can enrich trajectories, but it does not uniformly improve coherence and may shorten them. Model scale chiefly improves sustained generation, without producing reliable gains in canonical coherence or meaningful development. These results show that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities. WSE-bench makes those dynamics visible by extending narrative evaluation from finished stories to the processes that create them.
This work forms this challenge as Narrative Commitment Preservation (NCP), and introduces NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses that each environment includes a structured narrative specification that can automatically check throughout the interaction between the player agent...
Yingpeng Ma, Jianhao Yan, Bei-Ning Shi et al.· 1 citation
Vibe Narrativizing is formulated as turning natural-language writing requirements into a finished story, and MUSE, a Theory-Harnessed Story Engine, addresses two bottlenecks: rule quality and sustained rule realization.
Jian-Xiang Ma, Xiaocui Yang, Da-Ling Wang et al.· 0 citations
HappyWorld-Bench is introduced, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them, and highlights the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.
Zhi-Qi Bai, Ju-Nai Cai, Yi-Xin Chen et al.· 0 citations
Extensive experiments demonstrate that the CoDeR framework substantially extends the capabilities of existing world models, enabling long-term memory, open-ended interactions, autonomous evolution, and persistent multi-agent dynamics, while achieving state-of-the-art performance across multiple evaluation settings.
Zi-Xun Fang, Ya-Wen Shao, Kai Zhu et al.· 0 citations
CivBench is presented, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP), and two interface-level metrics are introduced that the environment makes measurable: Proactive Monitoring Rate (PMR) and RAG@10, capturing whether c...
Austin Andrews, L. Wilkinson, Jamie Heagerty et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.