Skip to content
Preprint

BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal

Aug 2026 · 1 citation · ⚡ 1 influential · 54 references
Computer Science

TL;DR

BeyondMasks reframes video object removal as causal scene consistency rather than local reconstruction and provides a unified framework for its evaluation, and proposes CORE, a structured vision language model based evaluation protocol that jointly measures object disappearance and after effect consistency, aligning more closely with human judgments than existing metrics.

Abstract

Recent advances in generative video models have significantly improved visual realism in video object removal, yet evaluation protocols still focus on masked region fidelity, treating removal as local inpainting. In real scenes, object removal is a causal intervention: eliminating an object also requires removing its induced physical effects, such as shadows, reflections, illumination changes, translucency, and dynamic traces. Existing benchmarks lack aligned clean references or remain limited to simplified synthetic settings, preventing systematic evaluation of causal consistency. We introduce BeyondMasks, a paired benchmark for causally consistent video object removal, consisting of temporally aligned synthetic and real world video pairs with clean background references. The dataset spans diverse photometric, geometric, volumetric, and dynamic interactions, and supports both mask based and instruction driven editing. We further propose CORE, a structured vision language model based evaluation protocol that jointly measures object disappearance and after effect consistency, aligning more closely with human judgments than existing metrics. Benchmarking state of the art methods reveals systematic failures in removing secondary physical effects despite high masked region fidelity, exposing a gap between visual plausibility and causal correctness. BeyondMasks reframes video object removal as causal scene consistency rather than local reconstruction and provides a unified framework for its evaluation.

View source

Similar papers

Preprint Aug 2026

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their gen...

Fei Wu, Wanke Xia, Xu He et al. · 0 citations
Preprint Sep 2026

Zero-Shot Object Removal via Attention Masking, Latent Anchoring, and Refinement

Removing an object from a real image requires more than synthesizing plausible content within a mask: the method must suppress residual object features, preserve the unedited scene, and generate replacement content that is consistent with the surrounding background. This paper approaches object removal from a stage-bas...

Arman Taghizadeh, U. Krumnack, Kai-Uwe Kühnberger · 0 citations
#artificial intelligence Preprint Aug 2026

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT is introduced, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation, and a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Disti...

Yu-She Cao, Shikun Feng, Ru-Xiang Duan et al. · 0 citations
Preprint Aug 2026

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

The OmniEdit-Bench provides a comprehensive and reliable testbed for evaluating instruction-based video editing and offers insights into future research directions, including an accuracy-aware penalty mechanism that conditions other scores on accuracy, preventing visually plausible but incorrect edits from receiving in...

Chenxuan Miao, Yutong Feng, Yi Lu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Beyond Visual Quality: Evaluating Physical Consistency under Ego-Motion with EgoGenEval

Recent visual generators produce high-fidelity images yet often violate physical consistency under ego-motion, limiting their use for spatial reasoning and embodied planning. Existing benchmarks largely focus on isolated images or single-step quality, leaving this challenge underexplored. We introduce EgoGenEval, a geo...

Yi-Lin Long, Chen-Ming Zhu, Zi-Tang Gou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

WorldCrafter is a video world model that learns a camera-queryable implicit 3D-aware memory that enables streaming scene exploration from a single input image or text prompt and shows substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploratio...

Wang-Bo Yu, Kunhao Liu, Wen-Bo Hu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.