Skip to content
Open access

TFCRNet: Dual-Discriminator SAR-to-Optical Translation and Region-Gated Cross-Attention Fusion for Thick-Cloud Removal

Sep 2026 · Remote Sensing · 0 citations · 45 references

Abstract

Thick-cloud contamination severely limits the usability of optical remote sensing imagery because cloud-covered regions may suffer from complete loss of surface information. Synthetic aperture radar (SAR) imagery provides complementary structural cues due to its cloud-penetrating capability, but the substantial cross-modal discrepancy between SAR and optical images makes high-fidelity SAR–optical fusion challenging. Existing methods usually either directly fuse heterogeneous SAR and optical features or use SAR-to-optical translation with insufficient spectral and structural constraints, which may lead to spectral distortion, structural artifacts, or degradation of cloud-free regions. To address these issues, we propose TFCRNet, a two-stage translation-and-fusion network for SAR–optical thick-cloud removal. In the translation stage, a Multi-Scale Feature Fusion Generator (MSFFG) transforms SAR imagery into optical-like images, while a Spectral Discriminator (SpeD) and a Structural Discriminator (StrD) separately constrain spectral fidelity and structural integrity. In the fusion stage, a Region-Gated Cross-Attention Fusion (RGCAF) module performs cloud-aware feature interaction between the translated optical image and the cloudy optical image. Using an externally supplied cloud mask, RGCAF emphasizes translated SAR-derived cues in cloud-covered regions while retaining reliable optical information in cloud-free regions. TFCRNet therefore requires a cloud mask during inference. Experiments on the SEN12MS-CR and SMILE-CR datasets show that TFCRNet achieves the best overall performance among the baseline methods reproduced under the unified experimental protocol adopted in this study. TFCRNet obtains 33.01/30.53 dB PSNR and 0.91/0.88 SSIM on SEN12MS-CR and SMILE-CR, respectively. These controlled results should be distinguished from literature-reported values obtained under different experimental settings, several of which are higher on selected metrics. Fine-grained ablations demonstrate that SpeD and StrD provide differentiated spectral and structural supervision, while RGCAF improves multimodal reconstruction through cloud-mask-guided regional information routing rather than spatially uniform feature fusion.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.