MMF-Net: Multimodal Mutual Feedback Fusion Network for Referring Remote Sensing Image Segmentation
Abstract
Referring remote sensing image segmentation (RRSIS) commonly treats language as a fixed query that only modulates visual features. This open-loop design is brittle in aerial scenes containing repeated objects, weak appearance cues, and relational expressions: visual evidence cannot revise which words and relations should dominate the query. We present MMF-Net, a closed-loop fusion architecture that couples stagewise reciprocal token refinement with decoder-wide context propagation. Its vision–language mutual feedback fusion (VLMFF) module updates visual and linguistic tokens through symmetric cross-attention and gated residual fusion, while multimodal context-aware fusion (MCAF) consolidates the corefined state and broadcasts it across the visual hierarchy. Under an identical Swin-B+ Bidirectional Encoder Representations from Transformer (BERT) backbone, decoder, training schedule, and data split, MMF-Net improves a strong internal baseline by 2.87/2.51/1.48 points at Pr@0.5/0.7/0.9 and by 0.68 oIoU and 0.58 mIoU. A factorial ablation further exposes positive VLMFF–MCAF interaction effects of 5.54 points at Pr@0.7 and 2.40 mIoU. MMF-Net reaches 64.86% mIoU and 78.14% oIoU on RRSIS-D with 115.2M parameters and 68.7 GFLOPs.