Scene-Level Planning for Temporally Coherent AI Video Generation Using Multimodal Representations
Recent text-to-video systems can generate visually appealing clips from natural language prompts, yet narrative prompts often contain multiple implicit temporal stages that require the generator to infer scene decomposition, subject persistence, action ordering, and visual continuity from a single unstructured input. T...