Remote Sensing Image Captioning via Dual-Stream Fusion and Spatial Relation-Aware Encoding
Experiments on the RSICD and NWPU-Captions datasets demonstrate that DSRAT achieves state-of-the-art performance across six metrics on RSICD and all seven metrics on NWPU-Captions, validating the effectiveness of the proposed approach.