Skip to content
Conference

Cross-Modal Dynamic Aggregation with Adaptive Relevance Modulation Fusion Network for Remote Sensing Visual Question Answering

2026 · Poster Volume 0007 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada · 0 citations

Abstract

Remote Sensing Visual Question Answering (RS VQA) task aims to provide accurate answers to questions about RS images. However, the semantic gap between low-level visual features and high-level semantics complicates the understanding of complex questions. Moreover, the lack of dynamic modulation mechanisms for integrating visual and textual features impedes the balanced interpretation of image content and textual semantics. To this end, we propose the Cross-Modal Dynamic Aggregation with Adaptive Relevance Modulation Fusion Network for RS VQA (CDAR-Net). Specifically, we propose a Cross-Modal Dynamic Aggregation Module (CMDA) that employs a multi-level attention mechanism to iteratively fuse text-guided visual features, enabling a progressive transition and dynamic integration from low-level visual features to high-level semantic information. We further introduce an Adaptive Correlation Modulation Fusion Module (ACMF) that dynamically adjusts visual and textual feature weights based on questions context, enhancing the representation of relevant information. Experimental results demonstrate that CDAR-Net outperforms existing state-of-the-art methods on three RS VQA datasets.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.