World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). In VDM-based WAMs, observations are first encoded by a VAE, and the resulting compressed latents are then processed by large video diffusion backbones to extract effectiv...
Xue-Ji Fang, Bo-Qiang Duan, Hua Wu et al.· 0 citations
Large-scale vision-language pretrained models like CLIP, paired with parameter-efficient fine-tuning techniques, have emerged as promising solutions for image-to-video transfer in video action recognition. However, existing methods often prioritize strong supervised performance at the expense of transferability and gen...
Meng-Meng Wang, Ze-Yi Huang, Bo-Yuan Jiang et al.· IEEE Transactions on Pattern...· 0 citations
Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense frames into a single VLM context incurs \textit{context rot} and high cost. Existing video agents often rely on query-agnostic offline prepro...
Sen Yang, Bo-Qiang Duan, Jing Yang et al.· 1 citation
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In this paper, we build...
Wei-Hao Bo, Shan Zhang, Yanpeng Sun et al.· 1 citation· ⚡1
Video object removal is a fundamental yet challenging task in video editing. Despite recent progress, existing methods typically fall into two categories. Traditional approaches based on optical flow or attention mechanisms often introduce noticeable artifacts and yield unnatural results. In contrast, diffusion-based m...
Zi-Zhao Chen, Ping Wei, Guang Dai et al.· 1 citation
GroundShot is presented, a training-free, model-agnostic agentic framework for entity-grounded multi-shot generation that improves multi-shot consistency over existing methods while requiring no additional training or model modification.
Yixuan Lai, Tianjia Shao, Kun Zhou et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.