UniMamba: A Unified Cross-Modal Mamba Framework for Remote Sensing Semantic Segmentation
Abstract
Fusing optical imagery with complementary modalities (X-modality), such as light detection and ranging (LiDAR) and synthetic aperture radar (SAR), is essential for robust semantic segmentation in complex environments. Although recent modality-agnostic models improve generalizability beyond fixed-pair methods, they still lack a unified and efficient architecture to process diverse combinations, including RGB-X, multispectral (MSI-X), and hyperspectral (HSI-X). To address this, we propose UniMamba, a unified cross-modal fusion Mamba (FuseMa) framework for general multimodal semantic segmentation. Specifically, UniMamba employs a unified encoder that processes diverse 2-D–3-D inputs without modality-specific modifications. The encoder integrates HyperMamba blocks to jointly model spatial, spectral, and frequency features while capturing global context with linear complexity. Each encoding stage further incorporates a cross HyperMamba (CroHMa) module for explicit cross-modal interaction, seamlessly followed by a FuseMa module for complementary feature fusion. Extensive experiments on three benchmarks demonstrate that UniMamba outperforms state-of-the-art methods and generalizes robustly across diverse multimodal settings. Code is available at https://github.com/xumzhang.