BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal
BeyondMasks reframes video object removal as causal scene consistency rather than local reconstruction and provides a unified framework for its evaluation, and proposes CORE, a structured vision language model based evaluation protocol that jointly measures object disappearance and after effect consistency, aligning mo...