Skip to content
Preprint

Spatial Attention Noise Masking for Causally Sufficient Interpretability

Aug 2026 · 0 citations · 33 references
Computer Science

TL;DR

Quantitative evaluations demonstrate mask faithfulness, near-baseline classification performance across five classification tasks despite substantial masking of image information, and robustness to distribution shifts such as background swapping and natural adversarial examples.

Abstract

We present a novel causal approach to interpretability for computer vision models that dynamically masks the input image prior to classification. The interpretability of deep learning predictions is critical in high-stakes fields such as medical imaging, security, and autonomous driving. Most interpretability methods are applied passively to already trained models, which typically result in correlational rather than causal explanations. Existing causal interpretability methods are limited to post hoc analysis, weakening the causal claims. Additionally, existing active methods generally lack explanations that explicitly assign responsibility to input features. This work proposes a spatial attention noise masking framework that provides causal explanations about the features sufficient for the prediction. The proposed framework consists of: 1) a UNet-style mask generator, and 2) a Resnet18 encoder and linear classifier that classifies both masked and unmasked versions of an input image. The generated masks are regularized to be sparse and spatially smooth, while masked image embeddings are constrained to remain consistent with embeddings from the corresponding unmasked images. The resulting masks can be interpreted as feature attribution maps that are competitive with related interpretability methods while additionally providing strong causal explanations of model predictions. Quantitative evaluations demonstrate mask faithfulness, near-baseline classification performance across five classification tasks despite substantial masking of image information, and robustness to distribution shifts such as background swapping and natural adversarial examples. Qualitative comparisons further demonstrate mask behavior and competitive interpretability relative to state-of-the-art feature attribution methods.

View source

Similar papers

Preprint Aug 2026

PatchGen: Learning Soft Intra-Image Predictive Subsets for Visual Generalization

Visual classifiers are expected to generalize under data shifts, target shifts, and their combinations, yet most existing methods focus on domain invariance while failing to address intra-image predictive sufficiency. We investigate the structural hypothesis that each image contains a sample-adaptive oracle intra-image...

Zhaorui Tan, Weimiao Yu, Xi Yang · 0 citations
Open access Sep 2026

Causal deconfounding for robust link prediction in multimodal knowledge networks

Link prediction in multimodal knowledge networks jointly exploits topology, text, and images, but visual backgrounds may introduce spurious correlations that reduce robustness under distribution shift. We propose MKGC-CSR, a causal deconfounding framework that models visual context as a confounder and approximates back...

Bing-Hong Li, Chu-Xin Xiao · 0 citations

Scalable Perturbation-Based Explanations via Tradeoff-Conditioned Smooth Masking

This work introduces a sampling-free, perturbation-based training framework based on continuous and differentiable masking that achieves competitive or superior attribution faithfulness compared to strong sampling-based baselines, while dramatically reducing computational cost and enabling substantially improved scalab...

Mehdi Naouar, Jens Rahnfeld, Yannick Vogt et al. · 0 citations
Open access Sep 2026

SPURIOUS CORRELATIONS IN EXPLAINABLE AI: DETECTION, QUANTIFICATION AND MITIGATION

Modern explainable AI (XAI) methods are widely used to audit deep neural networks, yet a growing body of evidence shows that the very explanations they produce can be hijacked by spurious correlations - statistical shortcuts between non-causal input features (image background, demographic attributes, textual artefacts)...

O. Verbytskyi · 0 citations
#machine learning Preprint Sep 2026

JEPA Learns What the Mask Leaves Unrecoverable

Joint-embedding predictive architectures are unusually sensitive to how the input is masked: block masks work, scattered masks do not, and the explanations are empirical. We give a measurement account. A mask is a linear measurement, and in a compactly supported wavelet basis every atom whose support lies inside the hi...

Peng Xie, Amr Alanwar · 0 citations
Preprint Sep 2026

Look Before You Judge: Training-Free Region Mining for Grounded and Explainable Deepfake Detection

Look Before You Judge, a training-free framework that formulates explainable deepfake detection as a sequential evidence acquisition process, and improves detection accuracy by up to 12.8%, reduces CHAIR by up to 33.4% and hallucination rate by up to 21.3%, and outperforms representative training-free decoding and atte...

Chia-Ling Chen, Yu-Ting Ta, Jian-Yu Jiang-Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.