Skip to content
Open access

Remote Sensing Image Captioning via Dual-Stream Fusion and Spatial Relation-Aware Encoding

Jul 2026 · The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences · Vol XLIX-B3-2026, pp. 41-47 · 0 citations · 13 references

TL;DR

Experiments on the RSICD and NWPU-Captions datasets demonstrate that DSRAT achieves state-of-the-art performance across six metrics on RSICD and all seven metrics on NWPU-Captions, validating the effectiveness of the proposed approach.

Abstract

Abstract. Remote sensing image captioning (RSIC) aims to describe key objects in remote sensing images using natural language, with significant applications in disaster assessment, land-use identification, and scene understanding. Existing methods face two critical challenges: insufficient cross-modal alignment due to the domain gap between generic visual representations and remote sensing semantics, and inadequate spatial relation modeling among regions in complex scenes, which compromises the semantic precision and logical coherence of generated descriptions. To address these issues, this paper proposes the Dual-Stream Relation-Aware Transformer (DSRAT) for remote sensing image captioning. On the visual encoding side, multi-scale CNN features serve as the foundation, fused with domain-specific semantic priors from RemoteCLIP through a gated dual-stream fusion module to achieve adaptive alignment of multi-source visual information. Subsequently, a spatial relation-aware mechanism is introduced into the encoder self-attention, which explicitly encodes geometric relationships such as relative position, distance, and orientation between regions as attention biases, enhancing the model’s capability for structured representation of complex spatial layouts and multi-object interaction scenarios. Finally, adaptive weighted aggregation of multi-layer encoder outputs generates discriminative cross-modal memory representations for the decoder. Experiments on the RSICD and NWPU-Captions datasets demonstrate that DSRAT achieves state-of-the-art performance across six metrics on RSICD and all seven metrics on NWPU-Captions. In particular, DSRAT achieves a significant performance improvement of +14.45 CIDEr on NWPU-Captions compared to the state-of-the-art method, validating the effectiveness of the proposed approach.

Read PDF

Similar papers

Open access 2026

GeoVP: A Unified Visual Prompting Framework for Multisource Remote-Sensing Image Understanding

GeoVP is proposed as a visual prompting MLLM for multisource RS image understanding, enabling unified image-level and region-level understanding under different prompt granularities and demonstrating the effectiveness of explicit visual prompting for prompt-conditioned region understanding in multisource RS imagery.

Le Yu, Yuan-Wen Wang, Xiao-Tong Qi · 0 citations
Open access Aug 2026

Soft Attention Enhanced CNN and LSTM Based Framework for Semantic Description Generation of Remote Sensing Imagery

Remote sensing images are complex, which makes it difficult to interpret and generate semantically appropriate textual description. To get a semantically relevant description, it is important to identify complex objects and understand the contextual relationships between them. In such cases, deriving contextually accur...

D. Pawade, Sonali Patil, R. Arya et al. · 0 citations
Open access Sep 2026

RS-CARES: Context-Aware Cross-Modal Alignment with Semantic Spatial Prior for Referring Remote Sensing Image Segmentation

Remote sensing referring image segmentation faces critical challenges including arbitrary target rotation, drastic scale variation, cluttered complex backgrounds, and large visual-language semantic gaps. Existing mainstream segmentation models adopt fixed-receptive-field backbones, coarse unidirectional cross-modal int...

Hui Xiong, Wen Luo, Bing He · 0 citations
2026

MMF-Net: Multimodal Mutual Feedback Fusion Network for Referring Remote Sensing Image Segmentation

Referring remote sensing image segmentation (RRSIS) commonly treats language as a fixed query that only modulates visual features. This open-loop design is brittle in aerial scenes containing repeated objects, weak appearance cues, and relational expressions: visual evidence cannot revise which words and relations shou...

Chongyang Li, Chen Wang, Wen-Kai Zhang et al. · 0 citations
Open access Sep 2026

Semantic-Guided Fusion Network for Multisource Remote Sensing Image Classification

Multisource remote sensing image classification has attracted increasing attention due to the complementary spectral, structural, and geometric information. However, existing methods still suffer from two limitations: insufficient semantic contextual modeling and unreliable feature fusion caused by slight spatial misal...

Yu-Wei Zhao, Chuan-Zheng Gong, Bao-Gui Huan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.