Skip to content

A Patch-Masking Approach to Explain CNNs and Vision Transformers through Class-Specific Impacts

· 0 citations · 9 references

TL;DR

A new method is presented for explaining CNNs and ViTs image classification through a set of interpretable rules composed of one or more antecedents, which combine pixel-level properties and patch-level features, analyzing the impact of each region on the model’s classification.

View source

Similar papers

Open access Aug 2026

Adaptive Feature Integration in CNN–Transformer Networks for Efficient and Interpretable Visual Classification

A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.

Komal Sharma, Monika Sainger · 0 citations
Open access Sep 2026

Vision Transformer Architectures for Next-Generation Image Classification: Attention-Driven Visual Understanding

Vision Transformer (ViT) architectures have emerged as powerful alternatives to conventional convolutional neural networks for image classification because they model long-range visual dependencies through self-attention. This paper presents a software-based image classification framework that uses a pre-trained ViT-Ba...

Sadeqa and Dr. Bitla Prabhakar · 0 citations
Open access Sep 2026

A data efficient pyramid vision transformer for image classification

A novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism is introduced.

Gazi Jannatul Ferdous, Medhi Hasan Chowdhury, Md. Azad Hossain et al. · 0 citations
Aug 2026

Comparative Study of CNN, Hybrid, and Transformer Architectures in Medical Image Classification.

The results show that larger models and larger pretraining datasets do not automatically lead to better downstream performance, and transfer effectiveness in medical imaging is driven primarily by architectural inductive biases, pretraining strategy, and domain relevance.

Dina A. Elkholy, Mohamed S. Shehata, John W. Braun · 0 citations
Sep 2026

An Explainable Deep Learning Framework for Robust Image Classification and Semantic Understanding

This paper evaluates a compact, from-scratch convolutional neural network (CNN) on the MNIST handwritten-digit benchmark along four simultaneous axes: classification accuracy against classical baselines, explainability faithfulness, adversarial and random-noise robustness, and feature-space separability. The CNN is imp...

I. Sengol · 0 citations
Review Aug 2026

Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations

Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or conc...

AmirHossein Eshghi, Hamid Saadatfar, S. A. Hoseini et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.