Skip to content

AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report

Jul 2026 · arXiv.org · Vol abs/2607.18367 · 4 citations · 34 references
Computer Science

TL;DR

AlayaWorld is presented, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p and introduces a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk.

Abstract

Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response. We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p. Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts. Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning. To reduce long-term drift, the model is trained with corrupted histories and prediction residuals collected from its own roll-outs. We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk. On iWorld-Bench, AlayaWorld achieves the best performance over long-horizon generation. Conceived as a full-stack, open-source, and long-term project, AlayaWorld is intended to provide an extensible foundation for future research on interactive video world models.

View source

Similar papers

Jul 2026

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld enables open-ended real-time interaction, allowing users to freely navigate and perform diverse actions such as combat, spell casting, and monster summoning, and the framework unifies the complete development-from data preparation model architecture, model training, inference acceleration, and deployment-within a modular and extensible architecture.

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li et al. · 1 citation
Preprint Aug 2026

Sekai2: From World Exploration to Interactive World Modeling

Sekai2 is introduced, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling, and Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant temporal descriptions.

Kang He, Wenshuo Peng, Zihui Gao et al. · 1 citation
Preprint Aug 2026

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

StreamMind is introduced, a two-tier architecture that assigns latency-critical interaction and proactive monitoring to independently scheduled frontend workers, while backend workers asynchronously construct persistent multimodal memory and perform historical recall and external search.

Xichen Zhang, Guankai Li, Yinghao Zhu et al. · 0 citations
Preprint Aug 2026

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

This work presents LiveAnimate, to their knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT).

Yuxuan Zhang, H. Xiong, Yubo Huang et al. · 0 citations
Preprint Aug 2026

AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

The new version of AlayaWorld substantially revise how conditioning signals are represented and integrated into the model, replacing the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer.

AlayaWorld Team Kaipeng Zhang, Chuanhao Li, Y. Zhan et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.