A Decoupled Dual-Teacher Collaborative Semisupervised Framework for High-Resolution Remote Sensing Image Semantic Segmentation
Abstract
Semantic segmentation of high-resolution remote sensing images is a fundamental task in Earth observation, yet its performance is often constrained by the expensive and time-consuming process of pixel-level annotation. While existing semisupervised learning (SSL) methods can mitigate this label scarcity, they often overlook the inherent spectral–spatial duality of remote sensing data. This oversight leads to the entanglement of heterogeneous features and the progressive accumulation of pseudolabel noise, ultimately limiting segmentation accuracy. To address this, we propose a decoupled dual-teacher collaborative semisupervised framework (DDCSF), which generates high-quality supervisory signals from heterogeneous modalities via novel decoupling and alignment mechanisms. DDCSF employs separate spectral and spatial teacher networks to achieve decoupled learning of texture features and geometric structures, respectively. To align these heterogeneous features, we design a frequency-aware cross-modal fusion module (FCFM). This module leverages the high-frequency components of spatial features to refine target boundaries and the low-frequency components to correct regional semantics, producing collaborative features with both precise details and semantic consistency. Furthermore, a pixel confidence voting module (PCVM) is introduced to quantify prediction uncertainty from the dual teachers and fusion module, selecting high-confidence pseudolabels to guide student model training. The student model’s ability to represent multimodal inputs is further enhanced by an adaptive spectral–spatial image fusion module (ASSIFM). Experimental results on the ISPRS Potsdam and GID benchmark datasets demonstrate that DDCSF consistently outperforms state-of-the-art methods at various label rates, validating its effectiveness.