#artificial intelligence
May 2026
Bernini: Latent Semantic Planning for Video Diffusion
To better handle multiple visual inputs, Segment-Aware 3D Rotary Positional Embedding (SA-3D RoPE) is introduced, and chain-of-thought reasoning in the planner is incorporated to better transfer understanding into generation.
Chenchen Liu, Junying Chen, Lei Li et al.
· arXiv.org · 10 citations
· ⚡2