A simple fusion of two novel components of residual attentive information forms a robust volume of selective residual attentive patterns (named SRAP), which boosted the performance of lightweight CNN-based networks by up to ~7% on ImageNet-100 without increasing the computational complexity.
Abstract
Modern deep networks often rely on attention modules, which are still at a modest level due to using either one type of channel-wise pattern or an expensive combination of two types of them. In the case of using all of those, the obtained weights can be less discriminative due to the disjointed excitations, while the model complexity would double. To deal with these limitations, an efficient attention is proposed by addressing two novel components of residual attentive information as follows: 1) top- $n$ channel-residual attentive patterns with a unitary excitation perceptron, and 2) multiple spatial-residual attentive features. A simple fusion of these complementary components forms a robust volume of selective residual attentive patterns (named SRAP). Experiments on benchmark datasets for image classification have proved the prominent performance of SRAP versus other attention modules. Particularly, SRAP boosted the performance of lightweight CNN-based networks by up to ~7% on ImageNet-100 without increasing the computational complexity. The implementation code of SRAP is available at https://github.com/nttbdrk25/SRAP.
A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.
Komal Sharma, Monika Sainger· International journal of com...· 0 citations
A comparative study of different deep learning architectures, including classical CNNs, deep hierarchical models, residual and dense networks, and compound-scaled architectures is presented, showing that deeper networks provide better representation, while residual connections and compound scaling improve training stab...
Riyaz Mohammed· International Journal of App...· 0 citations
An enhanced image recognition model integrating an adaptive improved pooling module and parameterized activation functions (Xexp) is proposed that outperforms comparison algorithms in recognition accuracy and stability.
Deeper modern networks outperform the older AlexNet by a wide margin on CIFAR-10, and even a relatively compact ResNet can nearly match the accuracy of a much larger VGG16 in far less time.
A new method is presented for explaining CNNs and ViTs image classification through a set of interpretable rules composed of one or more antecedents, which combine pixel-level properties and patch-level features, analyzing the impact of each region on the model’s classification.
Jean-Marc Boutay, Damian Boquete, Deniz Köprülü et al.· 0 citations