Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 51 references
Computer Science
TL;DR
The Spectral-Masked Bidirectional Fusion Transformer (SMBFT), a dual-stream architecture for spatio–spectral representation learning, combines a lightweight 3D convolutional frontend with transformer-based global modeling and constructs parallel spectral and spatial token streams to preserve modality-specific information.
Abstract
Hyperspectral image (HSI) classification requires robust learning from high-dimensional spectral signatures and spatial context, but remains challenging due to inter-band redundancy, limited labeled samples, class imbalance, and spatial heterogeneity. This paper proposes the Spectral-Masked Bidirectional Fusion Transformer (SMBFT), a dual-stream architecture for spatio–spectral representation learning. SMBFT combines a lightweight 3D convolutional frontend with transformer-based global modeling and constructs parallel spectral and spatial token streams to preserve modality-specific information. A bidirectional fusion module with symmetric cross-attention enables mutual interaction between the two streams, while a learnable gate regulates spatial-to-spectral information transfer to reduce over-fusion. In addition, a training-only band-masked spectral reconstruction objective is introduced as a tunable regularizer for improving spectral representation learning under limited supervision. Experiments on Houston 2013, Pavia University, Salinas, and Indian Pines show that SMBFT achieves competitive accuracy under the conventional random pixel protocol. To address spatial leakage concerns in patch-based HSI evaluation, we further quantify train-test patch overlap and report a buffer-constrained Houston 2013 evaluation with zero train-test patch overlap, where SMBFT maintains strong performance. Additional shared-configuration, ablation, and efficiency analyses indicate that the proposed components provide complementary gains while preserving low computational cost.
Hyperspectral image classification (HSI) requires a model to distinguish subtle spectral differences while preserving the spatial structure of land-cover regions. CNN-based methods are effective for local spectral–spatial extraction, but their limited receptive fields can weaken broader context modelling. Transformer-b...
Deep learning-based hyperspectral image (HSI) change detection (HSI-CD) methods have achieved promising accuracy. However, HSIs contain hundreds of highly correlated bands, while existing models often encode all bands with equal priority, resulting in redundant spectral computation. At the same time, bi-temporal featur...
Yan-Heng Wang, Kai Qin, Zhuan-Feng Li et al.· IEEE Journal of Selected Top...· 0 citations
Hyperspectral images (HSIs) provide rich spectral information, offering unique advantages for fine-grained land-cover classification. However, HSI classification remains challenged by insufficient spectral–spatial feature exploitation and significant variations in class difficulty under limited labeled samples. To addr...
Jin Zhang, Ying Cui, Li-Guo Wang et al.· IEEE Geoscience and Remote S...· 0 citations
Hyperspectral images (HSIs) possess fine spectral resolution. They can capture continuous and detailed spectral curves of ground objects, providing rich information for accurate classification. However, real-world scenes commonly suffer from diverse ground object morphology, spectral variability, and insufficient spati...
Shu-Fang Xu, Wei-Wen Xu, Shu-Yu Fei et al.· IEEE Transactions on Geoscie...· 0 citations
Hyperspectral image (HSI) classification under imbalanced and small-sample conditions is often hindered by the neglect of spatial contextual dependencies and the high-dimensional spectral redundancy. Moreover, conventional deep learning methods are prone to majority-class bias under imbalanced data distributions, which...
Wei Feng, Yan Cao, Yi-Jun Long et al.· IEEE Geoscience and Remote S...· 0 citations
Hyperspectral image (HSI) classification has been widely applied in numerous fields. Although deep learning-based methods have improved classification performance, existing approaches still struggle to balance accuracy and computational efficiency. Convolutional neural network (CNN)-based methods are limited by local r...
Han-Zhong Li, Hua Huang, Hong-Feng Li et al.· IEEE Transactions on Geoscie...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.