A unified collaborative dual-task learning framework that jointly optimizes S2O image translation and semantic segmentation through a shared hierarchical Vision Transformer is proposed, which achieves competitive S2O translation quality and semantic segmentation performance.
Abstract
Synthetic Aperture Radar (SAR) images have all-weather, day-and-night observation capabilities. However, compared with optical images, their speckle noise and non-intuitive scattering mechanism limit the interpretability of the images. Generative models for SAR-to-optical (S2O) conversion can improve visual interpretability, but existing methods often ignore the constraints on semantic structure, which are necessary for downstream tasks, for the sake of visual effects. We propose a unified collaborative dual-task learning framework, termed BMT (Bridging Modalities and Tasks), that jointly optimizes S2O image translation and semantic segmentation through a shared hierarchical Vision Transformer. The framework integrates: (1) a LocalViTBlock that fuses global self-attention with spatial depthwise convolution through a learnable gating mechanism; (2) an enhanced output module combining multi-scale refinement processing, color correction and anti-aliasing, which calibrates channel-level color statistics through feature fusion; (3) a ControlNet-style conditional injection mechanism that encodes SAR wavelet features and segmentation labels into a multi-scale feature pyramid and injects them at each encoder layer through zero-initialized convolution; (4) a bounded Kendall uncertainty weighting scheme that prevents either task from dominating the shared representation. We evaluate the framework under both paired and unpaired translation settings, on the public WHU-OPT-SAR paired dataset and a self-constructed unpaired ship dataset built from HRSID and DIOR, respectively. The experimental results show that the proposed method achieves competitive S2O translation quality and semantic segmentation performance. The dataset and source code have been publicly released at https://github.com/Lewisyuaner/BMT-S2O-main.
Fusing optical imagery with complementary modalities (X-modality), such as light detection and ranging (LiDAR) and synthetic aperture radar (SAR), is essential for robust semantic segmentation in complex environments. Although recent modality-agnostic models improve generalizability beyond fixed-pair methods, they stil...
Xu-Ming Zhang, N. Yokoya, Xing-Fa Gu et al.· IEEE Transactions on Geoscie...· 0 citations
Optical imagery provides rich spectral and texture cues but is vulnerable to cloud cover and imaging conditions, whereas synthetic aperture radar (SAR) offers all-weather observation but contains speckle noise and geometry-dependent distortions. Existing optical–SAR segmentation methods often treat local spatial correc...
Hao-Tian Liu· International Conference on...· 0 citations
A progressive layer-wise contrastive learning framework (MTC-Net) that couples the pseudo-label with the network’s representational hierarchy, forming a curriculum from local texture robustness to global semantic invariance.
Xiao Xiao, Han Zhang, Kenan Cheng et al.· Remote Sensing· 0 citations
Experiments on the Nanjing and public SEN1-2 datasets demonstrate that DSAN outperforms state-of-the-art models—including Pix2PixHD, CycleGAN, MSTMNet, and ICMA—in perceptual distribution realism (FID) with the sharpest geometric boundaries.
The proposed fusion-boundary-aligned routing regulates each modality's contribution before the first learned cross-modal feature-value mixing operation, supported by Spearman correlations between the learned routing weights and model-specific leave-one-modality-out utility range from 0.45 to 0.66.
Haifan Zhang, Yijing Wang, Haoyu Wang et al.· 0 citations
Synthetic aperture radar (SAR) has become an indispensable tool in Earth observation due to its capability for all-weather and day-and-night data acquisition. However, unlike optical sensors, the inherent coherent imaging mechanism of SAR results in single-channel grayscale images lacking intuitive spectral information...
Yong-Kang Chen, Peng Wang, Yu-Hang Xiao et al.· IEEE Transactions on Geoscie...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.