Skip to content
Open access

C3Former: A Spectral–Spatial Fusion Transformer for Hyperspectral Landcover Classification

2026 · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · Vol 19, pp. 25663-25687 · 0 citations · 45 references

Abstract

Hyperspectral remote sensing imagery provides rich spectral–spatial information for land-cover classification, offering advantages in complex scene understanding. However, existing methods often fail to fully exploit spectral–spatial dependencies due to limited labeled samples, leading to degraded performance under low-label-ratio settings. To address this issue, a novel hyperspectral classification model, termed C3Former, is proposed to explicitly model spectral–spatial dependencies and enable multiscale feature fusion through adaptive receptive field modulation. Specifically, a dynamic hole rate grouping atrous spatial pyramid pooling module is introduced to adaptively adjust receptive fields for objects at varying scales. By leveraging deformable sampling and multirate dilated convolutions, it effectively fuses multiscale spatial contexts and enhances feature discriminability. In addition, a spectral window cross-band attention transformer module is designed to perform frequency-aware spectral decomposition and hierarchical attention modeling, capturing local spectral continuity via window-based self-attention and modeling global cross-band dependencies through a cross-band attention mechanism. Furthermore, a cross-layer convolutional-transformer encoder module is developed to reinforce spectral–spatial dependencies via cross-layer feature interaction and gated fusion. Experiments on six public datasets demonstrate that the proposed method achieves superior performance. Extensive experimental results further validate its effectiveness, robustness, and generalization under diverse conditions.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.