Skip to content
Preprint

ConCA: Concentration-Aware Channel Attention for Fine-Grained Visual Recognition

Aug 2026 · 0 citations · 28 references
Physics Computer Science

TL;DR

ConCA is proposed, which pairs the mean with a shift-invariant negative-input entropy (NegEnt), computed via a softmax over the negated activations, forming a dual descriptor that jointly encodes magnitude and concentration.

Abstract

Lightweight channel attention mechanisms are widely used in image classification, yet their effectiveness in fine-grained visual recognition (FGVR) remains limited. Most modules summarize each channel by global average pooling (GAP), which captures activation magnitude but ignores spatial concentration, so channels with different spatial distributions but identical means receive the same descriptor. We propose Concentration-Aware Channel Attention (ConCA), which pairs the mean with a shift-invariant negative-input entropy (NegEnt), computed via a softmax over the negated activations, forming a dual descriptor that jointly encodes magnitude and concentration. A depthwise 1-D convolutional multi-layer perceptron (MLP), whose parameter count is linear in the number of channels, maps the pair to a per-channel weight. On six fine-grained benchmarks, ConCA improves over attention-free, SE-Net, and ECA-Net baselines as well as four richer descriptor-based modules under a controlled from-scratch protocol, and it generalizes across eight backbones on iNat2021-mini. These results indicate that the channel descriptor, together with the per-channel gating that maps it to attention weights, is an important but underexplored aspect of lightweight channel attention in FGVR.

View source

Similar papers

Open access Sep 2026

LGF-Net: a spatial-frequency deepfake detection network with learnable gabor filters and gated cross-modal Fusion

Deepfakes have become increasingly realistic due to recent advances in face manipulation techniques, making reliable detection in unconstrained environments more challenging. Existing spatial-frequency deepfake detection methods often rely on fixed hand-crafted frequency transforms and simple fusion strategies, which m...

Mukesh Pandey, Deepika Koundal, Sanjeev Kumar · 0 citations
Open access 2026

Threshold-Tuning Coordinate Attention for Binarized Neural Networks

—Binary Neural Networks (BNNs) are highly attractive for mobile and embedded vision due to their extremely low memory footprint and efficient bit-level convolutions. However, binarization often causes severe information loss and weakens the effectiveness of conventional attention modules designed for full-precision net...

Shao-Qing Wu, Hiroyuki Yamauchi · 0 citations
Preprint Sep 2026

Extending Decoupled Attention to Dense Prediction and Masked Training for Multi-Channel Images

Multi-Channel imaging (MCI) data differs fundamentally from natural images, as each channel records a semantically distinct signal rather than a colour band. To adapt vision encoders to MCI data, Multi-Channel Vision Transformers (MC-ViTs) tokenize each channel independently and concatenate the resulting tokens into on...

Umar Marikkar, S. Husain, Muhammad Awais et al. · 0 citations
Preprint Sep 2026

MSCA-UNet: Multi-Scale Context and Attention U-Net for Image Segmentation

U-Net remains a practical baseline for image segmentation because of its simple encoder-decoder structure and skip connections. However, the bottleneck representation is still dominated by a limited set of receptive fields, while decoder features are propagated without explicitly emphasizing the most informative channe...

Sheng-Wei Chan · 0 citations
Review Open access 2026

A Review of Fine-Grained Visual Categorization with Deep Learning

: Fine-grained visual categorization (FGVC) presents a class of recognition problems in which the discriminative signal is spatially concentrated, visually subtle, and easily destroyed by the preprocessing and augmentation strategies that serve coarse recognition well. Where standard image classification requires a mod...

Richard Adusei, G. Abdul-Salaam · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.