Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits that must preserve the original scene dynamics. Video scene text editing replaces text on scene surfaces, such as storefront signs, whiteboards, and product labels, while pr...
Xing-Hao Chen, Xiang-Bo Gao, Jiong-Ze Yu et al.· 0 citations
Monocular depth estimation foundation models, such as the Depth Anything series, have achieved remarkable performance across diverse domains. However, they still suffer from critical failures under adverse weather conditions, such as fog, rain, snow, or at night. To address this, we present Weather-Conditioned Depth An...
Zhao Xu, Chan-Wei Hu, Kuan Huang et al.· 0 citations
Visko Orbis 1.0 achieves the best DOVER aesthetic and technical scores and the best VideoAlign visual and motion quality, and leads three physical-plausibility protocols (VideoPhy-2, Physics-IQ, and VBench-2.0 Physics); in long-form Arena comparisons, it obtains the highest overall-preference and temporal-stability rat...
Xiang-Bo Gao, Siyuan Yang, Ping He et al.· arXiv.org· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.