Skip to content
Open access

A Deep Hybrid CNN–Transformer Framework Combining EfficientNet and Vision Transformer for Multiclass Skin Cancer Classification

Jul 2026 · Dandao Xuebao/Journal of Ballistics · Vol 38, pp. 216-223 · 0 citations

TL;DR

The results demonstrate that the hybrid Efficient Net–ViT architecture provides a robust, scalable, and reliable solution for automated skin cancer diagnosis and establishes a foundation for clinical AI applications.

Abstract

Skin cancer remains one of the most prevalent and life-threatening dermatological diseases worldwide. Early and precise detection plays a vital role in improving patient survival rates and reducing treatment costs. This paper presents a hybrid deep learning framework that integrates EfficientNetB0 and Vision Transformer (ViT) architectures to perform multiclass classification of dermoscopic skin lesions. The model is trained on the HAM10000 dataset, which includes eight types of skin cancer lesions, using transfer learning and data augmentation to improve generalization. EfficientNetB0 efficiently captures local spatial and texture features, while ViT models global contextual dependencies through self-attention mechanisms. Experimental evaluation demonstrates that the hybrid model achieves a validation accuracy of 82.73%, outperforming EfficientNetB0 (80.25%) and ViT (81.12%) by 2.48% and 1.61%, respectively. Additionally, the proposed framework achieves a macro precision of 0.7512, macro recall of 0.6158, and macro F1-score of 0.6505, confirming its superior classification capability. These results demonstrate that the hybrid Efficient Net–ViT architecture provides a robust, scalable, and reliable solution for automated skin cancer diagnosis and establishes a foundation for clinical AI applications.

Read PDF

Similar papers

Conference Jul 2026

Attention-Guided Ensemble Deep Learning Framework for Automated Skin Cancer Classification

Skin cancer is one of the most common malignancies in the world and early and accurate dermoscopic diagnosis is crucial for better survival outcomes of patients. There are various limitations in current single model convolutional and transformer models, such as limited ability to capture local texture, multi-scale morp...

Dharavath Nagesh, Erukonda Jairam · 0 citations
Open access Jul 2026

HybridSkinLes: An Explainable CNN-and Transformer-Based Multi-Class Skin Lesion Classification Framework

Skin cancer is considered a deadly disease globally, and the timely identification of the disease may save human life. This research presents a CNN–Transformer-based fusion framework for automated multi-class skin lesion classification. This approach combines ResNet50 and Vision Transformer (ViT) to categorize skin les...

M. Aldossary, Hina Gull · 0 citations
Conference Open access 2026

A Novel Hybrid Deep Learning Architecture Combining CNNs, Vision Transformers, and Multi-Scale Attention for Enhanced Detection of Breast Cancer in Histopathology Images

The proposed HrybridViT-CAM, a hybrid deep learning system that integrates convolution neural networks, Vision Transformers, and multi-scale attention system in order to classify breast cancer using histopathology images was able to detect the malignant regions of interest (ROIs) like nuclei pleomorphism, atypia chroma...

S. Angayarkanni, Mithila R, Koushik Rithik et al. · 0 citations
Open access Sep 2026

Hybrid CNN–BiLSTM with Multi-Transformer Stacking for Skin Lesion Classification

Skin cancer is one of the most common diseases worldwide and, if left untreated, it can be life threatening. In this work, we propose a dual-branch deep learning framework that integrates a CNN–BiLSTM module for local spatial–sequential feature modeling with multiple transformer models (ViT, DeiT, SwinV2, and BEiT) for...

Maryem Zahid, Mohammed Rziza, Rachid Alaoui · 0 citations
Aug 2026

A Multi-head Attention Transformer with Convolutional Projections for Skin Cancer Image Classification

A hybrid approach that integrates convolutions with linear projections is introduced, combining the ability of convolutions to effectively capture local spatial features with the capacity of linear projections to model global relationships, leading to a more expressive and robust feature representation.

Marwa Kahia, Bassem Bsir, Fathi Kallel · 0 citations
Open access Jul 2026

An Explainable Hybrid Deep Learning Framework Integrating SepConv2D, DenseNet121, and Vision Transformer for Automated Skin Lesion Classification

Early and accurate diagnosis of skin lesions is essential for reducing the mortality associated with skin cancer and improving patient outcomes. Although deep learning has significantly advanced automated dermatological diagnosis, existing approaches often struggle to simultaneously achieve high classification accuracy...

Adesh V. Panchal, Manish M. Patel · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.