Skip to content
Open access

A Vision Transformer with Dynamic Masking and Cross-Modal Semantic Learning for Remote Sensing Scene Classification

Aug 2026 · Remote Sensing · 0 citations · 33 references

Abstract

Remote sensing scene classification plays a vital role in various Earth observation applications. Although supervised learning remains the dominant paradigm, vast quantities of unlabeled imagery remain significantly underutilized. To leverage these unlabeled resources and enhance categorization accuracy, we propose a novel framework based on the Vision Transformer (ViT) that integrates a dynamic masking strategy with a cross-modal semantic learning mechanism. Specifically, a dynamic masking strategy guided by a smooth reconstruction loss is designed to learn robust feature representations from unlabeled samples prior to downstream fine-tuning. Furthermore, we incorporate cross-modal learning to enrich semantic information, thereby addressing the inherent supervisory limitations of conventional one-hot labels. Comprehensive experiments demonstrate that the proposed method significantly improves classification accuracy while maintaining high pre-training efficiency and strong generalization capabilities.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.