Skip to content
Open access

RL-Axial: Learning When to Compute Global Context for Multimodal Segmentation

Oct 2026 · Cognitive Computation · Vol 18 · 0 citations · 55 references

Abstract

High-resolution remote sensing semantic segmentation is a key enabler for fine-grained urban mapping and related geospatial applications. However, complex textures, severe scale variations, and cross-modal discrepancies make models with fixed architectures and computation paths struggle to combine reliable multimodal fusion with adequate global-context modeling. To address these challenges, we propose RL-Axial, a multimodal segmentation framework with interpretable, sample-adaptive context aggregation. A cross-modal gated fusion head (CMGF) performs pixel-wise fusion between RGB and digital surface model (DSM) features across multiple scales. A controllable axial-attention operator then provides four alternatives—off, row, col, and row+col—instead of applying a single axial configuration to every image patch. We train a lightweight selector with a group-normalized policy-gradient objective based on per-sample relative rewards. During the second training stage, the modality encoders and CMGF are frozen, whereas the policy, axial operator, decoder, and segmentation head are updated through their respective policy and segmentation losses. On the evaluated ISPRS Vaihingen and Potsdam splits, a single run of RL-Axial obtains 84.58 and 86.73 mIoU, respectively, numerically 0.35 and 0.53 points above the literature-reported FTransUNet values. The source-specific baseline protocols preclude a controlled superiority claim. Qualitative examples also show cleaner boundaries and more coherent object structures.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.