Comparative Performance Analysis of Vision Transformer (ViT) and Convolutional Neural Network (CNN) Architectures for Semarang Batik Motif Classification
Jul 2026· International Conference on Information and Communicatiaon Technology· pp. 1-6· 0 citations· 14 references
Abstract
Vision Transformers (ViT) capture global image context through self-attention but are data-hungry, typically underperforming Convolutional Neural Networks (CNNs) on the small datasets common in fine-grained tasks such as batik motif recognition. This study investigates whether a ViT, trained via knowledge distillation using the Data-efficient Image Transformer (DeiT), can overcome this limitation and compete with CNNs on a small Semarang Batik dataset. A distilled DeiT-Tiny student learns from a ResNet-50 CNN teacher and is benchmarked against two CNN references: ResNet-50 itself (a substantially larger model) and EfficientNet-B0 (a parameter-matched counterpart). In establishing this comparison, we first uncover a critical dataset integrity issue: the publicly available Semarang Batik Dataset (3,020 images) originates from only 18 unique source photographs, each augmented approximately 167 times prior to publication. This near-duplication causes severe data leakage under conventional random splitting, inflating the test accuracy of all models to a misleading 100% and rendering such evaluation meaningless. We therefore introduce a source-aware splitting strategy that enforces group-level separation between training, test partitions, and evaluate all models across three random seeds for statistical reliability. Under this corrected protocol, the distilled DeiT-Tiny attains the highest mean accuracy (95.18 ± 0.30%) and the lowest variance among the three models, matching both the larger ResNet-50 (94.87%) and the parameter-matched EfficientNet-B0 (94.68%) while using only 5.5M parameters. These results confirm knowledge distillation enables a compact Vision Transformer to compete CNNs on a limited fine-grained dataset, and underscore that verifying sample independence is a prerequisite for trustworthy evaluation on pre-augmented public datasets.
A novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism is introduced.
Gazi Jannatul Ferdous, Medhi Hasan Chowdhury, Md. Azad Hossain et al.· Discover Artificial Intellig...· 0 citations
A new method is presented for explaining CNNs and ViTs image classification through a set of interpretable rules composed of one or more antecedents, which combine pixel-level properties and patch-level features, analyzing the impact of each region on the model’s classification.
Jean-Marc Boutay, Damian Boquete, Deniz Köprülü et al.· 0 citations
Convolutional neural networks (CNNs) have long played a central role in computer vision due to their fast computation speed and high image feature extraction performance. CNNs can effectively learn diverse visual features ranging from low-level to high-level representations, resulting in high computational efficiency....
The results show that larger models and larger pretraining datasets do not automatically lead to better downstream performance, and transfer effectiveness in medical imaging is driven primarily by architectural inductive biases, pretraining strategy, and domain relevance.
Dina A. Elkholy, Mohamed S. Shehata, John W. Braun· Journal of imaging informati...· 0 citations
A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.
Komal Sharma, Monika Sainger· International journal of com...· 0 citations
Batik is an Indonesian cultural heritage featuring a wide variety of motifs with high inter-class visual similarity. The visual and manual identification process is highly subjective and inefficient. To address this classification problem, this study implements a deep learning approach using a hybrid Convolutional Neur...
Kevin Novebrianto, Hadi Zakaria· Jurnal Indonesia : Manajemen...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.