Skip to content
Preprint

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

This paper presents FADE, an effective training framework for counterfactual discovery and explanation that is built on an evidence-first, two-stage training paradigm, and demonstrates remarkable robustness when transitioning from constrained MCQs to unconstrained OQA and captioning.

Abstract

Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ) benchmarks inadvertently leak target events through their questions and candidate options. This reduces the core challenge from active discovery to text-guided verification. In this paper, we present FADE, an effective training framework for counterfactual discovery and explanation. Our method is built on an evidence-first, two-stage training paradigm. First, evidence-internalized supervised fine-tuning grounds the model's predictions in decisive visual anomalies. Second, we apply a fading-anchor reinforcement learning strategy that progressively removes textual guidance, compelling the model to independently discover and explain evidence. To rigorously evaluate this capability, we also introduce an effective pipeline that converts existing MCQ datasets into aligned MCQ, open-ended question answering (OQA), and captioning tasks without requiring additional data curation. Our simple approach yields strong results. Using Qwen3-VL-8B as the baseline, FADE achieves state-of-the-art strict paired scores across all three tasks on DualityVidQA-test and IPV-Bench, outperforming GPT-5.6. In specific, when transitioning from constrained MCQs to unconstrained OQA and captioning, our model demonstrates remarkable robustness. Its performance retention is 90.4% and 67.4% on DualityVidQA-test-substantially higher than the 48.1% and 30.7% retained by GPT-5.6. We hope this simple framework can serve as a solid baseline for future research in unconstrained counterfactual video understanding.

View source

Similar papers

#small language model Preprint Sep 2026

Counterfactual Reasoning for Robust Visual Question Answering

A novel training framework that enhances counterfactual contrastive learning for VQA and introduces a three-stage curriculum for stable multi-objective optimization, an enhanced Batch-Contrastive loss for more discriminative feature learning and two novel regularizers.

Truong-Binh Duong, T. Tran, Ngoc-Thao Nguyen et al. · 0 citations
#computer vision Preprint Sep 2026

From Knowing to Abstaining: Bridging the Representation-Action Gap in Vision-Language Models

The ability of vision-language models (VLMs) to abstain from unanswerable questions is as important as their ability to answer answerable ones accurately. Recently, several benchmarks have emerged to evaluate and improve VLM abstention, but they have substantial limitations. First, samples often contain shortcut cues i...

Jia-Luo He, Huang-Xun Chen · 0 citations
Preprint Aug 2026

REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering

ReVEAL consistently outperforms both closed-source and open-source state-of-the-art methods across extensive experiments and shows that explicitly verifying evidence sufficiency, rather than stopping at semantic relevance, retrieves the decisive clues that prior methods miss and yields more reliable long-video reasonin...

C.J. Yan, Yang Zhou, Meixing Shi et al. · 1 citation
Preprint Aug 2026

Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding

Clue-OPSD, a clue-privileged on-policy self-distillation framework for long-video understanding that uses clue intervals as privileged supervision without relying on ground-truth answer labels, while requiring no clue annotations or additional modules at inference time is introduced.

Kaishen Wang, Dong-Di Zhao, Yijun Liang et al. · 4 citations
Preprint Sep 2026

Counterfactual Attention Policy Distillation for Temporal Video Grounding

Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of On-policy distilla...

Shao-Bo Ju, Hai-Yang Yu, Xue-Cheng Wu et al. · 0 citations
Preprint Aug 2026

Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

Consistent gains over DVD on LVBench, Video-MME, EgoSchema, and LongVideoBench suggest that option-aware evidence acquisition transfers beyond MMR-V, and proposes PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video evidence acquisition.

Baixuan Xu, Yinyui Xu, Tianshi ZHENG et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.