Preprint
Jul 2026
PE-Field 4D: Video Generation Models as Canvas
This work revisits the role of positional encoding in video diffusion transformers and shows that it provides a useful spatial bias for geometry-aware control, and introduces a geometry-aware cross-attention mechanism that enables target video latent tokens to attend to structured context tokens derived from reference images or frames.
Yunpeng Bai, Haoxiang Li, Qi-Xing Huang
· 0 citations