This paper presents FADE, an effective training framework for counterfactual discovery and explanation that is built on an evidence-first, two-stage training paradigm, and demonstrates remarkable robustness when transitioning from constrained MCQs to unconstrained OQA and captioning.
Fufangchen Zhao, Jin-Hu Fu, Jiachen Lei et al.· 0 citations
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generato...
Jia-Shu Zhu, Yan-Hao Zheng, Rui Tian et al.· 0 citations
PhyParam is presented, a physics-guided image-to-video diffusion model that conditions on object-level forces, masses, friction, restitution, and scene-level gravity via a lightweight physical-attention routing mechanism, and further improves motion learning with semantic-structural feature-space supervision.
Yan-Xun Li, Hao Wen, Bingze Song et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.