When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations
WSE-bench is introduced, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling, showing that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities.