Skip to content
Open access

Cross-domain remote sensing scene classification via Swin Transformer with domain-invariant feature learning

Sep 2026 · Discover Computing · Vol 29 · 0 citations · 29 references

Abstract

Remote sensing scene classification is essential in various Earth observation tasks. However, its accuracy often declines when images are acquired from different sensors or modalities due to domain shift. The differences in imaging mechanisms, sensor properties, and data distributions create significant discrepancies between the source and target domains, making cross-domain generalization challenging. Though domain adaptation methods have presented encouraging outcomes in reducing these drawbacks, the challenges of efficient feature separation, sensor heterogeneity, and distributional mismatch in multimodal datasets remain unresolved. To overcome these disadvantages, the present paper proposes a hierarchical and context-aware feature extraction framework based on the Swin Transformer for cross-domain remote sensing image classification. The proposed approach employs the Swin Transformer to capture both local and global spatial dependencies, enabling better domain-invariant feature representation across heterogeneous sensor modalities. The proposed framework is evaluated on the Multimodal Remote Sensing Scene Classification (MRSSC2.0) dataset across three difficult cross-domain transfer tasks: VIS→SWI, VIS→INF, and VIS→SAR. A thorough comparison among ten domain adaptation methods is conducted to determine the effectiveness of the proposed feature extraction approach. Experimental results show that the proposed framework is particularly effective when domain differences are high, such as in the VIS→SAR transfer. Further evaluation demonstrates the effectiveness of the proposed feature extraction framework when combined with existing domain adaptation methods, with FixBi achieving 96% accuracy on the VIS→SWI task and CAN achieving 85% accuracy on the VIS→INF task. For the more challenging VIS→SAR transfer task, the proposed CMFA further improves cross-modal feature alignment, increasing the overall accuracy from 70% to 73% for FixBi and from 60% to 63% for CAN. These results demonstrate the effectiveness of the proposed approach in reducing the optical–radar modality gap while maintaining robust cross-domain classification performance.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.