Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 task instances with 2-1...
Jia-Cheng Hua, Xiao-Kun Feng, Jia-Qi Hua et al.· 0 citations
The architectural improvements and novel training recipe allow PaGE to achieve state-of-the-art performance on several gaze estimation tasks, outperforming humans in 7 out of 9 metrics while reducing the human-AI gap by at least 60% in the remaining 2.
Zhou-Tong Ye, Cheng-Wen Zhang, Zhai-Bin Cui et al.· arXiv.org· 0 citations
U-Lens improves verification efficiency and effort allocation, reduced perceived workload, and strengthened support across all three stages of uncertainty management, and reframes uncertainty support for generative AI from text-centered cues to a user-centered process of interpreting, evaluating, and acting on uncertai...
Yu Mei, Qingyue Zhuang, Jie Cai et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.