Lightweight optical–SAR semantic segmentation via joint modality–offset attention
Abstract
Optical imagery provides rich spectral and texture cues but is vulnerable to cloud cover and imaging conditions, whereas synthetic aperture radar (SAR) offers all-weather observation but contains speckle noise and geometry-dependent distortions. Existing optical–SAR segmentation methods often treat local spatial correction, modality weighting, and feature fusion as separate operations, leaving local correspondence evidence disconnected from modality selection. Our core contribution is Joint Modality–Offset Attention (JMOA), implemented in JMOA-Net, which performs shared competitive normalization over joint modality–offset candidates so that local candidate retrieval and modality contribution are determined by the same correlation scores. Supporting class-balancing and center-detail refinements are incorporated in JMOA-Net-Opt without changing JMOA's joint candidate normalization. On YESeg-OPT-SAR, JMOA Net-Opt achieves 79.17% mIoU, 90.12% overall accuracy, and 80.24% Kappa. Removing any individual component reduces mIoU by 1.01–4.32 percentage points, with the largest drop after removing JMOA. These results identify joint modality–offset competition as the primary contributor to the performance of the complete model.