Skip to content
Open access

Adaptive Feature Integration in CNN–Transformer Networks for Efficient and Interpretable Visual Classification

Aug 2026 · International journal of computer information systems and industrial management applications · 0 citations

TL;DR

A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.

Abstract

Over the past few years, deep learning has changed substantially following the emergence of Transformer architectures, which are particularly effective for representing long-range dependencies that are difficult for conventional Convolutional Neural Networks (CNNs). Whereas CNNs are well suited to extracting local spatial features using convolutional operations, Transformers are effective at representing global context through self-attention. Hybrid CNN–Transformer architectures have been developed to combine the respective strengths of the two approaches. A limitation of many existing models is their reliance on static or manually designed fusion strategies, which can restrict adaptability, add computational cost, and make the resulting decisions harder to interpret. The present study develops a novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods. The resulting framework is intended to improve both computational efficiency and model interpretability, thereby addressing important limitations of current hybrid designs. The experimental evaluation uses benchmark datasets such as ImageNet, CIFAR-100, and medical imaging datasets. The reported results show that the proposed model performs better than the comparison architectures with respect to accuracy, efficiency, and interpretability.

Read PDF

Similar papers

Open access 2026

Boosting Lightweight CNN-Based Networks Via Selective Residual Attentive Patterns for Image Recognition

A simple fusion of two novel components of residual attentive information forms a robust volume of selective residual attentive patterns (named SRAP), which boosted the performance of lightweight CNN-based networks by up to ~7% on ImageNet-100 without increasing the computational complexity.

Thanh Tuan Nguyen, H. Pham, Thinh Le Vinh et al. · 0 citations
Open access Aug 2026

TransCat: a hybrid CNN-transformer network with KAN for medical image segmentation

TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.

Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al. · 0 citations
Open access Aug 2026

Lightweight Hybrid CNN-Transformer Architecture for Diabetic Retinopa-thy Grading from Fundus Images: A Feature Fusion Approach with Dual Explainability

Automated diabetic retinopathy grading using fundus images demands accurate, computationally efficient, robust, and interpretable models. Although convolutional neural networks are effective at extracting local lesion patterns, they have limited ability to model long-range spatial relationships across the retina, while...

Rasha Jamal Hindi · 0 citations
Aug 2026

Comparative Study of CNN, Hybrid, and Transformer Architectures in Medical Image Classification.

The results show that larger models and larger pretraining datasets do not automatically lead to better downstream performance, and transfer effectiveness in medical imaging is driven primarily by architectural inductive biases, pretraining strategy, and domain relevance.

Dina A. Elkholy, Mohamed S. Shehata, John W. Braun · 0 citations
Open access 2020

A Comparative Study of Deep Learning Architectures for Image Classification

A comparative study of different deep learning architectures, including classical CNNs, deep hierarchical models, residual and dense networks, and compound-scaled architectures is presented, showing that deeper networks provide better representation, while residual connections and compound scaling improve training stab...

Riyaz Mohammed · 0 citations
Conference Jul 2026

Comparative Performance Analysis of Vision Transformer (ViT) and Convolutional Neural Network (CNN) Architectures for Semarang Batik Motif Classification

Vision Transformers (ViT) capture global image context through self-attention but are data-hungry, typically underperforming Convolutional Neural Networks (CNNs) on the small datasets common in fine-grained tasks such as batik motif recognition. This study investigates whether a ViT, trained via knowledge distillation...

Rafi Alifa Bagja, B. Purnama · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.