DSDFNet: Dual-Stream Dense-Feedback Network for RGB-T Semantic Segmentation
Abstract
RGB-T multi-modal semantic segmentation has demonstrated immense potential in complex scene understanding. However, existing methods are prone to introducing background noise during cross-modal feature interaction. Furthermore, traditional cascaded decoders frequently suffer from the dilution of deep semantic information during progressive upsampling. To address these limitations, this paper proposes a highly parameter-efficient Dual-Stream Dense-Feedback Network (DSDFNet). First, built upon a pure convolutional backbone (ResNet152), we design a Gated Cross-Modal Fusion Module (CMFM). This module extracts channel attention via dual global pooling (GAP and GMP) and achieves the noise-free bidirectional injection of heterogeneous features through strong residual connections and a soft gating mechanism. Second, breaking away from the conventional layer-by-layer decoding paradigm, we propose a Dense Multi-scale Feedback Decoder (DFBD). This decoder adaptively concatenates and aggregates multi-level features from the encoder within a unified spatial scale, densely feeding deep semantic representations back into the high-resolution reference feature map. Extensive experiments on mainstream RGB-T datasets, including PST900 and MFNet, demonstrate that the proposed method outperforms existing state-of-the-art approaches in terms of mIoU and mAcc metrics, exhibiting exceptional robustness and generalization capabilities.