The multi-scale dual-path cascaded attention network (MDCAN) demonstrates competitive classification performance and a favorable trade-off between classification accuracy and model complexity.
Abstract
Skin cancer represents a major global health concern, where early detection is critical for improving patient outcomes. Despite considerable promise in dermoscopic image analysis, existing deep learning methods share three persistent limitations: single-stream architectures cannot simultaneously capture global lesion morphology and local fine-grained details; fixed-receptive-field convolutions fail to handle wide intra-class scale variation; and parallel attention modules lack a structured hierarchical refinement logic. To address these issues, we propose the multi-scale dual-path cascaded attention network (MDCAN). A dual-path architecture explicitly decouples global semantic and local detail representations, while a cascaded attention module (MSCA) enforces a sequential ‘channel → spatial → scale’ refinement pipeline within each path. Depthwise separable and dilated convolutions are incorporated to capture information across different receptive fields while limiting additional computational overhead. Evaluated on the ISIC2017, ISIC2019, and HAM10000 datasets over five independent runs, MDCAN achieved mean accuracies of 89.25 ± 0.22%, 91.82 ± 0.32%, and 93.81 ± 0.17%, respectively. With 17.49 M parameters, MDCAN demonstrates competitive classification performance and a favorable trade-off between classification accuracy and model complexity.
Existing multimodal approaches for breast cancer classification rely on fixed-stage fusion, where clinical features are incorporated as static auxiliary inputs, limiting dynamic interactions between visual and clinical information. Furthermore, reported performance gains in this literature are rarely validated under le...
Deema A. Alzamil, B. Alkhamees· IEEE Access· 0 citations
Accurate automated segmentation of melanoma from dermoscopic images remains a critical challenge in computer-aided
diagnosis pipelines due to high intra-class diversity, fuzzy lesion boundaries, and structural artifacts. Conventional deep
learning frameworks, including convolutional neural networks, vision transformers...
Sanjyoti Kumari Tarai, Renuka Arora· International Journal of Dru...· 0 citations
Multimodal skin-lesion classification requires integrating complementary clinical and dermoscopic images while addressing inter-branch disagreement and the semantic gap. We propose MDCL-SCRL, a modality-specific dual-stream framework using paired images and patient metadata. Multimodal Dynamic Consensus Learning (MDCL)...
Jin-Tao Liu, Jin-Nan Zhang, Tong Ling et al.· IEEE journal of biomedical a...· 0 citations
—The rapid ascent of deep learning in medical image analysis has been important in the remarkable advancements made in automated tumor identification. This study presents a novel convolutional neural network architecture for efficient tumor classification by integrating multi-scale fusion techniques with dual-path atte...
S. P. Praveen, Shaik Salma Begum, P. Padmavathi et al.· Journal of Advances in Infor...· 0 citations
Accurate lesion segmentation in ultrasound images remains challenging because of severe speckle noise, low tissue contrast, ambiguous boundaries, and substantial variations in lesion morphology and scale. To address these limitations, we propose RDPA-UNet, an enhanced U-shaped network that integrates differential-path...
Jia-Feng Jin, Hengsheng Zhang, Kun Wu· Mathematics· 0 citations
A novel Swin Transformer–based U-Net architecture whose primary contribution lies in a Swin-Enhanced Cross Attention (SECA)-driven decoding strategy, rather than the use of a Swin encoder alone, is proposed.
Reza Ahmadi Lashaki, Farhad Bayrami, Shayan Rokhva et al.· Multimedia tools and applica...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.