Video diffusion and flow models require many sequential evaluations, making generation computationally expensive. Few-step distillation reduces this cost but poses a capacity allocation problem: a student must match the teacher's iterative generation with far less sequential computation. Existing trajectory methods ask...
Lang Feng, Péter Karkus, M. Igl et al.· 0 citations
World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT reco...
Chensheng Peng, Wen-Hao Ding, Ran Tian et al.· 0 citations
This work proposes COMPACT-VA, a planning-aligned working memory framework built on conditional VQ-VAE, compressing extended context into bounded representations, and evaluates on high-signal dynamic scenarios where historical context is most critical for behavior correctness, and accordingly design behavioral metrics.
Zhixuan Liang, Yuxiao Chen, Yurong You et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.