Skip to content

Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

The analysis shows that CoTs do not reliably track visual evidence that influences model predictions, and it is found that Predict-then-Explain explanations align more strongly with perturbation-induced probability shifts than pre-answer CoTs, while binary vCT scores are often nearly saturated.

Abstract

Chain-of-thought (CoT) may often look plausible, yet it may not faithfully reflect the model's decision-making process. While methods for measuring the faithfulness of CoTs for textual inputs have been increasingly introduced, using these methods for visual inputs is not straightforward. In this work, we adapt the family of counterfactual methods for measuring CoT faithfulness, namely the Counterfactual Test (CT) and Correlational Counterfactual Test (CCT), to visual inputs, and call them vCT and vCCT, respectively. Using vCT and vCCT, we benchmark eight recent open-source Vision Language Models (VLMs) on two datasets. Our analysis shows that CoTs do not reliably track visual evidence that influences model predictions: they may omit the removed object even when its removal causes a large prediction shift, yet mention it when the shift is small. We further find that Predict-then-Explain explanations align more strongly with perturbation-induced probability shifts than pre-answer CoTs, while binary vCT scores are often nearly saturated. We also include a reconstruction control, in which images pass through the same editing pipeline without object removal, and find that the main object-removal intervention induces larger shifts than reconstruction alone. We construct and release Counter-SNLI-VE and Counter-A-OKVQA, two datasets of image pairs that differ by a single object.

View source

Similar papers

#large language models Open access Sep 2026

SRFE: Measuring Chain-of-Thought Faithfulness with Independent Stepwise Visual Probes

Using stepwise probe as an external reference or Chain-of-Thought (CoT) to evaluate visual-language model (VLM) analysis ability is common, but the faithfulness of PASS/FAIL verdicts in VLM produced by explicit CoT and independent SRFE probe needs further work to figure out. Prior works on CoT faithfulness are mainly a...

Xin-Yi Wei · 0 citations
#natural language process... Preprint Sep 2026

An Empirical Study of Counterfactual Self-Explanations in LLMs

This work evaluates ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales and shows that model scale is the strongest determinant of explanation quality.

Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis-Mastromichalakis et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness

Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through input--output behavior or input attributions, leaving internal computation largely underexplored. We instead cast faithfulness as internal concept grounding: Does a large lan...

Qianli Wang, Yilong Wang, Dennis Wei et al. · 0 citations
#machine learning Preprint Sep 2026

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

This work operationalizes the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts, and evaluates whether LLM judges can identify high-advantage steps and finds that sufficiently capable LLM...

Kevin Du, A. Hoyle, L. Ruis et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EDCT-Bench: Uncovering Faithfulness Gaps in VLMs via Explanation-Driven Counterfactual Testing

Vision-Language Models (VLMs) can produce Natural Language Explanations (NLEs) that sound plausible yet remain inconsistent with the visual evidence they cite. We present Explanation-Driven Counterfactual Testing (EDCT), an intervention-based protocol that extracts visual concepts cited in a model's explanation, applie...

Sihao Ding, Santosh Vasa, Aditi Ramadwar et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.