VideoArgus is introduced, a unified rubric-grounded framework covering five video generation and editing settings that achieves higher within-input Spearman and Kendall correlations with human judgments than the corresponding benchmark-specific evaluators across all five tasks.
Abstract
Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide limited evidence for their scores. We introduce VideoArgus, a unified rubric-grounded framework covering five video generation and editing settings. For each input instance, VideoArgus generates an output-blind, sample-specific rubric once and reuses it to evaluate all corresponding candidate videos. The rubric defines concrete criteria, scoring rules, failure modes, and evidence plans, which guide criterion-specific VLM QA and visual tools to produce evidence-grounded criterion scores, rationales, and a diagnostic report. We further construct VideoArgus-Bench, containing 1,026 curated input instances built from 653 high-quality images and 416 high-quality videos, with all benchmark rubrics pre-generated, frozen, and released. On a separate 1,260-video human-alignment set, VideoArgus achieves higher within-input Spearman and Kendall correlations with human judgments than the corresponding benchmark-specific evaluators across all five tasks. Model rankings also remain largely consistent across different rubric-generation and evaluation-VLM backbones. All code and data are released. Visit our project page: https://zzzmyyzeng.github.io/VideoArgus
VideoVIBE is introduced, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks and V2Lens is proposed, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-...
Jia-Jun Xu, Yang-Hao Zhou, Jing Liao et al.· 1 citation
AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from select...
Generative video systems can produce short clips from textual and visual instructions, yet their ability to preserve the content of a human reference remains uncertain. This exploratory study examines 160 videos generated with Doubao-Seedance-2.0 from 40 human-made short-form videos drawn from Douyin and Xiaohongshu. T...
Chloe Zihan Jin· Communications in Humanities...· 0 citations
Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. We introduce PhoenixNest-Video, an evidence-grounded multimodal agen...
Yuxuan Fan, Miao-Jun Huang, Hai-Mei Zhang et al.· 1 citation
Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 task instances with 2-1...
Jia-Cheng Hua, Xiao-Kun Feng, Jia-Qi Hua et al.· 0 citations
This work introduces VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants, and formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, an...
Fan Zhang, Guang-Ming Yao, Jin-Yang Wu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.