Skip to content
Open access

Orthogonal Feature Decoupling and Hierarchical Cross-Scale Aggregation network for RGB-T semantic segmentation

Sep 2026 · Measurement science and technology · 0 citations

Abstract

Visible-thermal (RGB-T) imaging systems provide crucial complementary information for robust visual sensing and measurement in complex illumination conditions. However, effectively fusing multi-modal sensory data to achieve accurate semantic segmentation remains a key challenge. Most existing methods rely on heuristic fusion strategies or standard attention mechanisms, which fail to decouple modality-shared semantics from modality-specific details leading to redundant representations. Furthermore, conventional decoders progressively upsample features in a stage-wise manner, which may introduce cross-scale inconsistencies due to mismatched spatial resolutions and receptive fields. To address these challenges in multi-sensor data processing, we propose ODCANet, an Orthogonal Feature Decoupling and Hierarchical Cross-Scale Aggregation Network, which consists of an Orthogonal Feature Decoupling Module (OFDM) and a Hierarchical Cross-Scale Aggregation Module (HCSAM). Specifically, OFDM employs Correlation-guided Decomposition Blocks (CDB) with global-local correlation modeling to estimate shared and modality-specific components, and applies an explicit orthogonality constraint during training to encourage lower correlation between them, thereby promoting more complementary fused representations. Meanwhile, HCSAM performs dynamic cross-scale aggregation of adjacent-stage features to help alleviate semantic gaps across decoder stages. Extensive experiments on benchmark datasets demonstrate that ODCANet achieves the best reported mIoU among the compared methods, reaching 62.21%, 89.25% and 67.98% mIoU on MFNet, PST900 and FMB datasets, respectively.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.