Jul 2026· International journal of computer information systems and industrial management applications· Vol 18, pp. 1-18· 0 citations
TL;DR
The findings simplify the selection of the model in clinical studies in implementing dermoscopy, suggesting that ViT-Small/16 should be used as a probabilistic-ranked malignancy screening model, and EfficientNetV2-S should be used as a fixed threshold binary triage model in a telemedicine system.
Abstract
Skin cancer is one of the life-threatening malignancies in the world, timely detection of which is directly proportional to the survival of patients. Traditional dermoscopic diagnosis suffers inter-observer variability, lack of specialists and is not scalable in resource limited healthcare environments. The end-to-end hierarchical feature learning provided by deep learning is a transformative solution to the learned dermoscopic image corpora. The paper empirically comparatively examines eight binary skin lesion classifiers benign versus malignant of a dataset of 2,637 training images, and 661 held out test images. The assessed architectures are located on a wide design range a custom CNN trained using fresh data, two ResNet18 transfer learning pipelines, DenseNet121, MobileNetV3-Large, ViT-Small/16 (ImageNet-21K), ConvNeXt-Tiny (ImageNet-12K) and EfficientNetV2-S (ImageNet-21K). All the models are trained with the same conditions involving stratified splitting, Weighted Random Sampler, two-stage fine tuning with discriminative learning rates, Automatic Mixed Precision, and early stopping. It is evaluated using six metrics accuracy, per-class precision, recall, F1-score, ROC-AUC, and PR-AUC. The highest test accuracy (91.53) and macro-F1 (0.907) is attained with EfficientNetV2-S. ViT-Small/16 has the best ROC-AUC (0.9723) and PR-AUC (0.9699), which proves the effectiveness of Vision Transformer in threshold-free probabilistic discrimination the clinically decisive measure in screening applications. The three contemporary timm-based models have consistently reached ROC-AUC 0.95 and above, but the legacy CNN models are at 0.56 even though the legacy CNN models are at competitive accuracy of 89-90%. MobileNetV3-Large yields a false negative rate of 61 (20% miss rate), which highlights the clinical risk of aggressive model compression. The findings simplify the selection of the model in clinical studies in implementing dermoscopy, suggesting that ViT-Small/16 should be used as a probabilistic-ranked malignancy screening model, and EfficientNetV2-S should be used as a fixed threshold binary triage model in a telemedicine system.
Skin diseases represent one of the most widespread categories of health disorders worldwide, and timely diagnosis plays a critical role in preventing complications such as skin cancer. Conventional diagnostic procedures depend largely on visual examination by dermatologists, a process that is subjective, time-consuming, and difficult to access in rural or under-resourced regions. This paper presents a comprehensive deep learning-based framework for the automated detection and classification of skin diseases from dermoscopic and clinical images. The proposed system employs a transfer-learning approach built on the ResNet50 convolutional neural network architecture, pre-trained on ImageNet and fine-tuned on benchmark dermatological datasets, namely HAM10000, the ISIC Archive, and DermNet. The methodology encompasses dataset collection, image pre-processing, data augmentation, feature extraction, model training, disease classification, and rigorous performance evaluation. ResNet50 is selected for its residual-learning capability, which mitigates the vanishing-gradient problem and enables deeper, more accurate networks suited to fine-grained medical image analysis. An extensive review of over thirty related studies spanning convolutional architectures, ensemble methods, and emerging transformer-based models is used to position the proposed framework within the current state of the art. The framework is evaluated using standard classification metrics, including accuracy, precision, recall, specificity, F1-score, Cohen’s kappa, and confusion-matrix analysis, and is benchmarked conceptually against alternative architectures such as a baseline CNN, VGG16, MobileNetV2, DenseNet, and EfficientNet. The anticipated outcome is an accurate, scalable, and accessible screening tool capable of assisting healthcare professionals in early diagnosis, thereby reducing diagnostic delay and improving healthcare accessibility, particularly in regions with limited dermatological expertise.
Nisha Rajodiya, Shailendra Mishra, Sumitra Menaria et al.· International Journal of Res...· 0 citations
GOALS
To compare a vision transformer with 2 convolutional neural network architectures for multiclass lesion classification in capsule endoscopy images.
BACKGROUND
Manual review of capsule endoscopy is time-consuming and subject to interobserver variability. Deep learning can automate lesion recognition; however, most prior capsule endoscopy work evaluates a small number of classes, and systematic comparisons between transformer and convolutional architectures across many lesion categories are limited.
STUDY
Two publicly available data sets (SEE-AI and Kvasir-Capsule) were merged and preprocessed to create a 21-class image data set (∼58,000 frames). Images were resized to 224 × 224 pixels and split using stratified sampling into training (n=40,587), validation (n=8696), and test (n=8696) sets. A pretrained Vision Transformer, DenseNet121, and ResNet50 were fine-tuned using categorical cross-entropy loss and Adam optimization, with early stopping. Performance was assessed using accuracy, macroaveraged precision, recall, F1 Score, and the area under the receiver operating characteristic curve.
RESULTS
On the independent test set, the vision transformer achieved 92.2% accuracy with macroaveraged precision/recall/F1-score of 0.92 and an area under the receiver operating characteristic curve of 0.99. DenseNet121 achieved 74.0% accuracy (F1-score 0.78; area under the receiver operating characteristic curve 0.85). ResNet50 achieved 38.0% accuracy (F1-score 0.40; area under the receiver operating characteristic curve 0.55).
CONCLUSIONS
In this merged 21-class capsule endoscopy image data set, the vision transformer achieved higher frame-level classification performance than DenseNet121 and ResNet50 under the present experimental conditions. Importantly, the data set was split at the image level rather than at the patient or procedure level, because frames from the same examination may be correlated; therefore, performance figures likely reflect benchmark results on this frame-level public data set and should not be interpreted as estimates of patient-level generalization or as evidence of definitive architectural superiority. These findings support further evaluation of transformer-based approaches, but grouped reanalysis, external validation, and workflow-oriented studies are required before clinical implementation.
S. Boppana, S. Komati, Aditya Chandrashekar et al.· Journal of Clinical Gastroen...· 0 citations
Skin lesion diagnosis remains one of the most challenging tasks in medical image analysis because several benign and malignant lesions exhibit overlapping visual characteristics such as irregular borders, non-uniform pigmentation, structural asymmetry, and texture similarity. Although deep learning methods have achieved remarkable progress in dermoscopic image classification, many existing systems depend exclusively on convolutional neural networks and may not fully benefit from clinically interpretable handcrafted descriptors or classifier confidence information. In addition, direct feature fusion strategies often fail to maximize the complementary relationship between handcrafted and deep representations.
This study proposes a Sequential Hybrid Meta-Learning Model for automated four-class skin lesion classification involving basal cell carcinoma, melanoma, nevus, and pigmented benign keratosis. The proposed framework integrates handcrafted dermatological image descriptors, pretrained deep convolutional neural network embeddings, and decision-level classifier confidence scores within a multi-stage learning pipeline. A confidence-margin data cleaning stage is introduced to reduce noisy or ambiguous samples, followed by dataset balancing and hierarchical cancer/non-cancer diagnostic decisioning.
Unlike conventional direct-fusion approaches, the proposed model learns sequentially by first extracting deep discriminative knowledge and then reusing classifier confidence outputs as meta-features for final decision making. Experimental evaluation demonstrates substantial performance gains over baseline hybrid systems. The model achieved 95.87% flat four-class accuracy, 94.11% hierarchical accuracy, 95.64% cancer/non-cancer Level-1 accuracy, 0.992245 ROC-AUC, and 0.991202 PR-AUC.
The results confirm that combining feature-level and decision-level knowledge significantly improves skin lesion discrimination and offers a practical computer-aided diagnostic solution for intelligent dermatology screening.
M. A. Belal, M. El-Gazzar, BenBella S. Tawfik et al.· Statistics, Optimization &am...· 0 citations
The results demonstrate that deep learning techniques can significantly assist in early detection and classification of skin cancer, thereby supporting dermatologists in clinical decision-making and improving diagnostic efficiency and mortality rates associated with skin cancer.
A. Star, Gibi Linza, Siva Durshika et al.· 0 citations
Skin cancer is a significant global health concern, where early and accurate detection is important for supporting effective treatment and improving patient outcomes. To perform manual examination of the skin lesions is time consuming and could rely significantly on clinical expertise, which may lead to the need of computer aided approaches helping to classify skin lesions. In this study, an automated, DL approach to classifying dermoscopic skin images as benign or malignant is proposed. In this research, a publicly available Kaggle skin cancer dataset was used, containing 3600 images with labels, half of which were malignant and the other half were benign. The data set was split into 70% training, 15% validation and 15% testing sets. The input images were preprocessed and enhanced using image processing and augmentation techniques to standardize images, enhance the diversity of training samples, and enhance model generalization. Transfer learning was explored in four different pretrained convolutional neural network (CNN) architectures. The models were trained under similar experimental conditions analysis were used for the evaluation. Of the evaluated architectures, MobileNetV2 had the best overall classification performance with an accuracy of 91.11%. The accuracy of DenseNet201 was 88.06%, higher than that of ResNet50V2 (86.81%) and Xception (85.83%). To explore the learning behavior of the models, training and validation accuracy and loss curves were also analyzed, in addition to the quantitative evaluation. Furthermore, an explainable artificial intelligence method named Grad-CAM was added to visualize the image regions that the model relied on to make its prediction, thus gaining another insight into the model's decision-making behavior of the CNN architectures. The results show that both the CNN architectures used for the experiment and the automated classification of skin lesions using them are effective, and the MobileNetV2 performs the best in the experimental framework of this study. The proposed framework can serve as a foundation for the development of efficient and interpretable computer-aided skin lesion classification systems.
Raees Adnan, Fawad Nasim, Muqaddas Salahuddin· SOCIAL PRISM· 0 citations
Skin cancer remains one of the most widespread and fatal malignancies globally, making early, accurate diagnosis essential for improving patient survival rates. While deep Convolutional Neural Networks (CNNs) have advanced automated dermoscopic image analysis, practical clinical adoption is hindered by severe dataset class imbalance, data leakage, and diagnostic opacity. To overcome these challenges, this study introduces STLC-Net (Skin Transfer Learning Classification Network), a robust, zero-data-leakage framework for automated seven-class skin lesion classification using the benchmark HAM10000 dataset. The methodology integrates in-graph dynamic data augmentation to eliminate framework shape-locking errors during deployment, and incorporates a log-smoothed class-weighting strategy to stabilize gradient descent against severe minority class scarcity. Under a strict zero-leakage evaluation protocol (2,003 isolated test samples), three distinct architectural paradigms—DenseNet121 (dense feature concatenation), EfficientNetB3 (compound scaling), and MobileNetV2 (inverted residuals)—were trained via a two-stage transfer learning protocol. DenseNet121 emerged as the superior standalone backbone, achieving an accuracy of 79.73%, precision of 83.02%, and an ROC-AUC of 0.9398. To further eliminate individual architectural blind spots, a soft-voting STLC-Ensemble was constructed, achieving peak performance with 80.38% accuracy, 83.53% weighted precision, 79.85% F1-score, and an ROC-AUC of 0.9492. The complete pipeline is integrated into an interactive web deployment for real-time, interpretable clinical decision support.