Skip to content
Review

Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review

Jul 2026 · arXiv.org · Vol abs/2607.09403 · 0 citations · 28 references
Computer Science

TL;DR

The architectural patterns validated here, including layer-as-budget compression, semantic-locality scheduling, and separation of generation and review, transfer to the broader class of knowledge-intensive, multi-agent LLM applications.

Abstract

Worldbuilding, the construction of coherent fictional worlds, is a foundational task in game design and literary creation. Large Language Models (LLMs) offer new possibilities for automated content generation, but their application to worldbuilding faces three challenges: context explosion that grows linearly with the building process, the tension between creative diversity and content consistency, and the absence of automated quality assurance. This paper presents AutoWorldBuilder, a multi-agent collaborative system that addresses these challenges through five integrated components: a structured concept network with conflict detection; a DAG-based hybrid batch scheduler that groups tasks by semantic locality; a four-layer context compression mechanism achieving approximately 90% token reduction; an iterative review system with specialized Auditor agents that improves proposal pass rates from 42% to over 85%; and a skill-driven agent architecture supporting zero-code extension with differentiated temperature configuration. Two experiments across 20 diverse worldbuilding tasks, using GPT-OSS 120B and DeepSeek v3.2 as LLM backends, demonstrate a 95.0% success rate. The system generated 56-103 self-consistent concepts per world in 18-31 minutes with zero-conflict delivery. The architectural patterns validated here, including layer-as-budget compression, semantic-locality scheduling, and separation of generation and review, transfer to the broader class of knowledge-intensive, multi-agent LLM applications.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

WorldBench is presented: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions, and Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and ot...

Leonardo Ranaldi, Sherrie Shen, Jushi Kai et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Exploring Collaboration between a language and a non-language agent

To solve LLM collaboration with non-language agents, latent state internalization is introduced, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state.

Harini S.I., Somesh Singh, Yaman Kumar Singla et al. · 0 citations
Open access 2026

Agentic GraphRAG and Deterministic Schema Reconciliation for High-Compliance Domains: An LLMOps and FinOps Approach

A scalable Agentic GraphRAG architecture structured under a comprehensive LLMOps Feature-Extraction-Inference (FTI) lifecycle, which achieves a 70% improvement in Citation Rate compared to standard Vector RAG and establishes a robust Safe Abstention rate, effectively mitigating the risk of ungrounded generation.

Marcelo Massashi Simonae, A. Ortoncelli, Marlon Marcon · 0 citations
Jul 2026

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents

PhoenixRepair is a multi-agent framework that systematically explores multiple candidate edit locations and performs iterative reflection and refinement on patch generation, thereby expanding the search space of repair strategies and achieves higher fault localization accuracy than existing approaches.

Tian-Yue Jiang, Yan-Lin Wang, Xin He et al. · 2 citations
Jul 2026

AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration

AgentRadio is presented, an asynchronous message-passing layer that equips coding-agent harnesses with three primitives: threads, messages, and waiting for mentions that shows the gain growing with task difficulty, consistent with mid-course correction as the underlying mechanism.

Xinxing Ren, Qianbo Zang, Ziyan Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.