Virtual try-on (VTON) requires precise pixel-level fidelity, yet mainstream Diffusion Transformers (DiTs) often suffer from texture degradation and structural drift. We identify symmetric interactions in standard joint-attention mechanisms as a source of these failures. Although such interactions support semantic flexi...
Zi-Shu Qin, Zhi-Yu Jin, Pi-Pei Huang et al.· 0 citations
InfinityEdit is proposed, a lightweight edit adapter that equips a streaming video generator with unbounded editing ability that faithfully continues the stream under each edit, and stays stable over unbounded edit sequences.
Yunze Tong, Mu-Shui Liu, Can-Yu Zhao et al.· 0 citations
This work formally identifies this limitation as High-Frequency Trajectory Collapse: supervised fine-tuning converges to the conditional mean of the training distribution, which is dominated by smooth, low-frequency textures, causing high-frequency patterns to become nearly un-sampleable.
Yuhan Li, Xianfeng Tan, Fan-Gao Zeng et al.· 0 citations
Self-OPD is introduced, a teacher-free OPD framework for flow matching models that turns the student's own self-exploration into step-wise supervision and outperforms prior RL and OPD methods without task-specific teachers.
Shi-Yi Zhang, Mu-Shui Liu, Yunze Tong et al.· 1 citation
This work proposes REST (Reward-Enhanced Scored-Trajectory Distillation), a single-stage RL-distillation co-training framework that attaches a decoupled student to an arbitrary RL teacher that enables few-step CFG-free inference that matches or surpasses its 40-step RL teacher, with an overall additional training cost...
Yuhan Li, Fan-Gao Zeng, Sicong Kang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.