Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The chall...
Xing-Ming Shui, Da-Peng Chen, Bo-Wei Liu et al.· 0 citations
This work introduces meta-detection into AI-generated video detection, enabling reliable forgery detection by jointly optimizing predicted labels and supporting evidence within reinforcement learning, and introduces evidence-aware credit assignment, which preserves reliable label supervision while encouraging detectors...
Bo-Wei Liu, Zheng Lu, Yuhan Bian et al.· 1 citation
Variance Recovery Policy Optimization (VRPO) is introduced, which retains and progressively expands groups to recover informative signals from prompts that are difficult yet solvable and retains and progressively expands these groups to recover informative signals from prompts that are difficult yet solvable.
Jingqi Tian, Haoji Zhang, Lin Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.