Review
Aug 2026
MLLM-Guided Semantic Correction for Text-to-Video Generation
A training-free, interpretable mid-generation correction framework that integrates multimodal large language model (MLLM) feedback directly into the diffusion sampling loop and achieves diffusion trajectory correction by injecting semantic evaluation signals during video synthesis, enabling the model to optimize the generated content through continuous self-reflection.
Junhao Chen, Zheqi Lv, Keting Yin et al.
· 0 citations