Vision Transformer (ViT) Based Image Recognition with Explainable AI & Self-Supervised Pretraining
Abstract
One recent architecture that offers an alternative to convolutional neural networks for picture recognition is the Vision Transformer (ViT). Compared to convolutional neural networks, these designs' use of self-attention allows them to better model global context. Inefficient training data and unintelligible decision-making are two basic obstacles that ViT models must surmount regardless of their effectiveness. These difficulties are of the utmost importance in systems that are safety-critical, have scarce labelled data, and have high model transparency. An image recognition framework based on ViT is presented in this research. To this, we have integrated explainable AI with self-supervised pretraining. The system lessens its reliance on labelled data by employing self-supervised representation learning. For accurate and clear predictions, it uses attention-based explainability. Extensive testing has shown that the proposed method improves visual explanations, data efficiency, recognition accuracy, and structural robustness. The outcomes prove that next-gen image recognition systems can be reliably developed using explainable ViT models that undergo self-supervised pretraining.