Preprint
Aug 2026
Beyond Text Conditioning: A Systematic Study of MLLM-DiT Fusion for Video Generation
BiVidGen, a hybrid framework where an MLLM first generates semantic visual tokens and a DiT renders videos conditioned on both text and these tokens via multi-layer cross-attention is proposed.
Yanbo Ding, Yijia Fan, Caihua Shan et al.
· 0 citations