Skip to content
Preprint

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing

Aug 2026 · 5 citations · 54 references
Computer Science

TL;DR

VideoArgus is introduced, a unified rubric-grounded framework covering five video generation and editing settings that achieves higher within-input Spearman and Kendall correlations with human judgments than the corresponding benchmark-specific evaluators across all five tasks.

Abstract

Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide limited evidence for their scores. We introduce VideoArgus, a unified rubric-grounded framework covering five video generation and editing settings. For each input instance, VideoArgus generates an output-blind, sample-specific rubric once and reuses it to evaluate all corresponding candidate videos. The rubric defines concrete criteria, scoring rules, failure modes, and evidence plans, which guide criterion-specific VLM QA and visual tools to produce evidence-grounded criterion scores, rationales, and a diagnostic report. We further construct VideoArgus-Bench, containing 1,026 curated input instances built from 653 high-quality images and 416 high-quality videos, with all benchmark rubrics pre-generated, frozen, and released. On a separate 1,260-video human-alignment set, VideoArgus achieves higher within-input Spearman and Kendall correlations with human judgments than the corresponding benchmark-specific evaluators across all five tasks. Model rankings also remain largely consistent across different rubric-generation and evaluation-VLM backbones. All code and data are released. Visit our project page: https://zzzmyyzeng.github.io/VideoArgus

View source

Similar papers

Preprint Aug 2026

VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

VideoVIBE is introduced, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks and V2Lens is proposed, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-...

Jia-Jun Xu, Yang-Hao Zhou, Jing Liao et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from select...

Gunin Gupta, Nirmit Arora, Pavan Tankala · 0 citations
#generative ai Review Sep 2026

Reference Fidelity in AI-Generated Short-Form Video Across Four Prompt Packages

Generative video systems can produce short clips from textual and visual instructions, yet their ability to preserve the content of a human reference remains uncertain. This exploratory study examines 160 videos generated with Doubao-Seedance-2.0 from 40 human-made short-form videos drawn from Douyin and Xiaohongshu. T...

Chloe Zihan Jin · 0 citations
#artificial intelligence Review Sep 2026

PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment

Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. We introduce PhoenixNest-Video, an evidence-grounded multimodal agen...

Yuxuan Fan, Miao-Jun Huang, Hai-Mei Zhang et al. · 1 citation
Preprint Sep 2026

ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts

Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 task instances with 2-1...

Jia-Cheng Hua, Xiao-Kun Feng, Jia-Qi Hua et al. · 0 citations
Preprint Aug 2026

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

This work introduces VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants, and formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, an...

Fan Zhang, Guang-Ming Yao, Jin-Yang Wu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.