Transformer-enhanced land use classification algorithm for high-resolution remote sensing imagery
Abstract
High-resolution remote sensing imagery provides rich spatial and semantic information for land use classification, which plays a crucial role in urban planning, resource management, and ecological monitoring. However, traditional convolutional neural network (CNN)-based approaches struggle to effectively capture long-range dependencies and complex contextual relationships inherent in such imagery. To address this limitation, we propose a Transformer-enhanced land use classification framework tailored for high-resolution remote sensing images. The method integrates local feature extraction via convolutional layers with global dependency modeling through Transformer attention mechanisms, thereby enabling both fine-grained texture recognition and contextual semantic understanding. Furthermore, a multi-scale feature fusion strategy and improved positional encoding are introduced to enhance representation robustness across varying resolutions. We evaluate the proposed model on four widely used benchmark datasets, including ISPRS Vaihingen, ISPRS Potsdam, UCM Land Use, and NWPU-RESISC45. Experimental results demonstrate that our approach achieves superior overall accuracy (OA), mean Intersection over Union (mIoU), and F1-scores compared to state-of-the-art baselines such as UNet, Deeplabv3+, and Swin Transformer. These findings validate the effectiveness of combining CNNs and Transformer mechanisms in advancing automatic land use recognition and provide a promising pathway for scalable applications in large-scale remote sensing analysis.