Skip to content
Conference

On the Instability of Saliency Maps under Sparse Perturbations

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 91-96 · 0 citations · 24 references

Abstract

Understanding the reliability of model explanations remains a critical challenge in deep learning. Prior work has shown that saliency maps can be manipulated by optimizing the input using gradient-based methods, where gradients of the loss with respect to the input are computed to generate dense perturbations that alter the model’s explanations. However, the impact of sparse perturbation constraints remains underexplored, largely due to the difficulty of optimization arising from their discrete and non-differentiable nature. In this work, we pioneer the investigation of this problem by employing a gradient-free sparse adversarial attack. Our goal is to modify only a small subset of pixels in an input image to successfully attack the model and induce distortions in the saliency maps, such that the highlighted important regions are displaced. Experimental results on both general-domain and medical vision models show that altering only a tiny fraction of pixels is sufficient to drastically change the model’s predictions and the spatial distribution of attention in the saliency maps. Our study provides new insights into the interplay between robustness and explainability against sparse perturbation, and highlights risks in relying on saliency-based explanation, particularly in safety-critical applications. Source code is available at https://github.com/ELO-Lab/SparseAttack-To-Saliency-Map.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.