Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfol...
Yu-Bo Zhu, Ya-Wen Shao, Zi-Yun Dai et al.· 0 citations
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and o...
Lianghua Huang, Zhigang Wu, Yupeng Shi et al.· arXiv.org· 2 citations
Wan-Streamer v0.2 is presented, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model, which raises the interactive output stream from 192x336 to 640x368 while preserving approximately 200 ms model-side signal-to-signal latency at 25 FPS.
Lianghua Huang, Zhigang Wu, Yupeng Shi et al.· arXiv.org· 2 citations
ReWorld separates the two during training and bounds them at inference, and under a three-axis protocol covering action following, long-horizon recall, and video quality, against six recent interactive world models it attains the best control fidelity.