A new method is presented for explaining CNNs and ViTs image classification through a set of interpretable rules composed of one or more antecedents, which combine pixel-level properties and patch-level features, analyzing the impact of each region on the model’s classification.
A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.
Komal Sharma, Monika Sainger· International journal of com...· 0 citations
Vision Transformer (ViT) architectures have emerged as powerful alternatives to conventional convolutional neural networks for image classification because they model long-range visual dependencies through self-attention. This paper presents a software-based image classification framework that uses a pre-trained ViT-Ba...
Sadeqa and Dr. Bitla Prabhakar· International Journal of Adv...· 0 citations
A novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism is introduced.
Gazi Jannatul Ferdous, Medhi Hasan Chowdhury, Md. Azad Hossain et al.· Discover Artificial Intellig...· 0 citations
The results show that larger models and larger pretraining datasets do not automatically lead to better downstream performance, and transfer effectiveness in medical imaging is driven primarily by architectural inductive biases, pretraining strategy, and domain relevance.
Dina A. Elkholy, Mohamed S. Shehata, John W. Braun· Journal of imaging informati...· 0 citations
This paper evaluates a compact, from-scratch convolutional neural network (CNN) on the MNIST handwritten-digit benchmark along four simultaneous axes: classification accuracy against classical baselines, explainability faithfulness, adversarial and random-noise robustness, and feature-space separability. The CNN is imp...
I. Sengol· Natural Resources for Human...· 0 citations
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or conc...
AmirHossein Eshghi, Hamid Saadatfar, S. A. Hoseini et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.