VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple objects and attributes. Existing diffusion-based enhancement methods fail to adequately address the unique challenge of cross-scale error propa...
Zhen-Nan Chen, Tianxing Shi, Pengcheng Xu et al.· 0 citations
Progress in general-purpose video editing depends on constructing large-scale paired supervision and effectively adapting video-generation backbones to instruction-driven editing. Unlike video generation, video editing must execute a requested transformation while preserving unrelated subjects, scene structure, motion,...
J.Jenny Li, Di Shao, Xin-Yu Chen et al.· 0 citations
InstructVVT is proposed, an instruction-driven and reference-guided video virtual try-on framework based on a Diffusion Transformer that operates without inference-time spatial priors that outperforms state-of-the-art open-source methods in garment fidelity, structural preservation, and temporal consistency, despite re...
Di Shao, Song-Han Wu, Xin-Yu Chen et al.· 0 citations
Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.