Skip to content

MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing

Sep 2026 · 0 citations
Computer Science

TL;DR

Vibe Narrativizing is formulated as turning natural-language writing requirements into a finished story, and MUSE, a Theory-Harnessed Story Engine, addresses two bottlenecks: rule quality and sustained rule realization.

Abstract

LLMs have been able to generate fluent prose, but high-quality stories also require coordinated decisions about plot, character, and language across planning, drafting, and revision. We formulate Vibe Narrativizing as turning natural-language writing requirements into a finished story. MUSE, a Theory-Harnessed Story Engine, addresses two bottlenecks: rule quality and sustained rule realization. Story theory supplies the rules, and a practical agent harness puts them to work. Knowledge engineering organizes Robert McKee's theory through rule atomization, semantic consolidation, mechanism abstraction, a single source of truth, and layered disclosure; typical examples clarify judgments that depend on context and aesthetic purpose. The harness preserves story decisions in intermediate deliverables across design, character performance, scene composition, and revision. Context engineering supplies each role with relevant guidance and decisions; a masterwork corpus provides inspiration and prose references. A worked example follows a requested object from its thematic role to climactic actions. Across four base models, MUSE improves WritingBench by 1.1 to 6.2 points over zero-shot generation; it is the only multi-stage system in our comparison to do so. It also raises LongStoryEval by more than ten points on three of the four models. ConStory-Bench consistency error density remains in the low single digits for all four models, below every reproduced story-system baseline on three of the four models. Ablations locate the largest quality contribution in structural design, voice-specific effects in the character path, and further gains in revision.

View source

Similar papers

Preprint Sep 2026

Incipit: Axiom-Grounded Scaffolding for Human-AI Literary Creation

Large language models can produce fluent prose from short prompts, but direct prompt-to-text interaction gives writers limited access to the assumptions that shape a long narrative. We present Incipit, an implemented research prototype that introduces an explicit planning layer between a writer's intent and generated p...

Qiang-Si Liu, Chun-Yi Zhao · 0 citations
Preprint Aug 2026

SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution

This work presents SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations, and releases PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes.

Mao-Lin Ran, Xiaoyan Lu, Jia-Qi Liu et al. · 0 citations
Preprint Sep 2026

What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation

Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coherently continue stories across image, narration, and speech....

Sieun Hyeon, Yejoo Lee, Mintaek Lim et al. · 0 citations
Open access Oct 2026

Do AI stories align with the laws of the story economy?

The article situates LLM-generated narratives within the twenty-first-century story economy by asking how they rework the storytelling patterns shaped by social media, algorithmic optimization, and the pursuit of the “compelling story.” Previous research has shown that stories gaining traction on social media conf...

Juha Raipola, Maria Mäkelä, Samuli Björninen et al. · 1 citation
Preprint Aug 2026

When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

WSE-bench is introduced, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling, showing that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities.

Yuqi Chen, Sixuan Li, Yunfeng Cai et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals, assumptions that poorly match how guidance is sought in re...

Yu-He Wu, Guang-Yu Wang, Yu-Jie Chen et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.