Skip to content
Preprint

Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

Sep 2026 · 0 citations
Computer Science

TL;DR

A unified collaborative dual-task learning framework that jointly optimizes S2O image translation and semantic segmentation through a shared hierarchical Vision Transformer is proposed, which achieves competitive S2O translation quality and semantic segmentation performance.

Abstract

Synthetic Aperture Radar (SAR) images have all-weather, day-and-night observation capabilities. However, compared with optical images, their speckle noise and non-intuitive scattering mechanism limit the interpretability of the images. Generative models for SAR-to-optical (S2O) conversion can improve visual interpretability, but existing methods often ignore the constraints on semantic structure, which are necessary for downstream tasks, for the sake of visual effects. We propose a unified collaborative dual-task learning framework, termed BMT (Bridging Modalities and Tasks), that jointly optimizes S2O image translation and semantic segmentation through a shared hierarchical Vision Transformer. The framework integrates: (1) a LocalViTBlock that fuses global self-attention with spatial depthwise convolution through a learnable gating mechanism; (2) an enhanced output module combining multi-scale refinement processing, color correction and anti-aliasing, which calibrates channel-level color statistics through feature fusion; (3) a ControlNet-style conditional injection mechanism that encodes SAR wavelet features and segmentation labels into a multi-scale feature pyramid and injects them at each encoder layer through zero-initialized convolution; (4) a bounded Kendall uncertainty weighting scheme that prevents either task from dominating the shared representation. We evaluate the framework under both paired and unpaired translation settings, on the public WHU-OPT-SAR paired dataset and a self-constructed unpaired ship dataset built from HRSID and DIOR, respectively. The experimental results show that the proposed method achieves competitive S2O translation quality and semantic segmentation performance. The dataset and source code have been publicly released at https://github.com/Lewisyuaner/BMT-S2O-main.

View source

Similar papers

2026

UniMamba: A Unified Cross-Modal Mamba Framework for Remote Sensing Semantic Segmentation

Fusing optical imagery with complementary modalities (X-modality), such as light detection and ranging (LiDAR) and synthetic aperture radar (SAR), is essential for robust semantic segmentation in complex environments. Although recent modality-agnostic models improve generalizability beyond fixed-pair methods, they stil...

Xu-Ming Zhang, N. Yokoya, Xing-Fa Gu et al. · 0 citations
Conference Sep 2026

Lightweight optical–SAR semantic segmentation via joint modality–offset attention

Optical imagery provides rich spectral and texture cues but is vulnerable to cloud cover and imaging conditions, whereas synthetic aperture radar (SAR) offers all-weather observation but contains speckle noise and geometry-dependent distortions. Existing optical–SAR segmentation methods often treat local spatial correc...

Hao-Tian Liu · 0 citations
Open access Aug 2026

MTC-Net: Leveraging Multi-Temporal Consistency and Multi-View Synergistic Contrastive Learning for Remote Sensing Scene Classification

A progressive layer-wise contrastive learning framework (MTC-Net) that couples the pseudo-label with the network’s representational hierarchy, forming a curriculum from local texture robustness to global semantic invariance.

Xiao Xiao, Han Zhang, Kenan Cheng et al. · 0 citations
Open access Sep 2026

DSAN: Dual-Scale Aligned Network with Asymmetric Priors and Differentiable Soft-Edge Loss for SAR-to-Optical Image Translation

Experiments on the Nanjing and public SEN1-2 datasets demonstrate that DSAN outperforms state-of-the-art models—including Pix2PixHD, CycleGAN, MSTMNet, and ICMA—in perceptual distribution realism (FID) with the sharpest geometric boundaries.

Ying-Ying Kong, Dong-Ming Wang · 0 citations
Preprint Aug 2026

Boundary-Aligned Contribution Routing for Robust Optical--SAR Object Detection

The proposed fusion-boundary-aligned routing regulates each modality's contribution before the first learned cross-modal feature-value mixing operation, supported by Spearman correlations between the learned routing weights and model-specific leave-one-modality-out utility range from 0.45 to 0.66.

Haifan Zhang, Yijing Wang, Haoyu Wang et al. · 0 citations
2026

PhDGAN: A Physics-Informed Dual-Branch GAN With Gamma-Prior for Dual-Polarization SAR Image Colorization

Synthetic aperture radar (SAR) has become an indispensable tool in Earth observation due to its capability for all-weather and day-and-night data acquisition. However, unlike optical sensors, the inherent coherent imaging mechanism of SAR results in single-channel grayscale images lacking intuitive spectral information...

Yong-Kang Chen, Peng Wang, Yu-Hang Xiao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.