2026· IEEE Transactions on Geoscience and Remote Sensing· Vol 64, pp. 5636413-5636413· 0 citations· 71 references
Abstract
Remote sensing visual question answering (RSVQA) remains challenging because the evidence relevant to the question in remote sensing (RS) imagery is typically sparse, spatially dispersed, and highly variable in scale and layout. This structural heterogeneity poses a major challenge to existing transfer strategies, which often struggle to achieve effective query-conditioned semantic focusing and localized spatial refinement. To address this issue, we propose CRISP, a task-specific parameter-efficient adaptation framework for RSVQA built on a frozen ViLT backbone. CRISP comprises two complementary components. First, a cross-modal residual guidance (CMRG) module generates instance-specific guidance tokens from pooled image and question summaries, steering early cross-modal interaction toward query-relevant content while suppressing background interference. Second, an attention-guided spatial realignment (ASR) module performs offset-guided feature realignment within intermediate Transformer layers, enabling localized refinement of spatial evidence under scale variation and sparse semantic distribution. Extensive experiments on the RSVQA-LR and RSVQA-HR benchmarks show that CRISP achieves strong overall performance and consistently improves overall accuracy (OA) and average accuracy (AA) over prior methods, with particularly notable gains on presence, comparison, and region-related questions. These results demonstrate that residual guidance and spatial realignment together provide an effective task-specific parameter-efficient adaptation strategy for RSVQA under the frozen-backbone setting. The code will be available at https://github.com/PhD-Xu/CRISP
CROSS is proposed, a tightly integrated paradigm for RRSIS that achieves state-of-the-art performance and maintains precise localization even under severe spatial description perturbations, standing as a robust new paradigm for RRSIS.
Tingzhang Luo, Ruizhong Liu, Yi-Chao Liu et al.· 1 citation
DARAD is proposed, a dual-adapter and ranking-aware distillation framework that preserves the historical cross-modal ranking structure while learning new visual and textual concepts from evolving archives to address the challenge of continual RS-ITR.
Xi Chen, Xu Chen, Xiang-Yang Jia et al.· 0 citations
Remote Sensing Visual Question Answering (RS VQA) task aims to provide accurate answers to questions about RS images. However, the semantic gap between low-level visual features and high-level semantics complicates the understanding of complex questions. Moreover, the lack of dynamic modulation mechanisms for integrati...
Zi-Hua Zuo· Poster Volume 0007 The 2026...· 0 citations
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...
Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al.· 0 citations
A selective tool use framework in which a single VLM either answers directly or invokes a deterministic change analysis tool to obtain question specific evidence to demonstrate the benefit of question-specific semantic evidence for Change VQA, while highlighting the influence of semantic prediction quality on the resul...
Y. Bazi, M. M. Al Rahhal, M. Mekhtiche et al.· 0 citations
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual question answering over satellite imagery, yet localizing specific objects or changed regions remains challenging. Existing approaches rely on generating bounding box coordinates as token sequences, which is fragile...
J. Chung, Sungjune Park, Yeongyun Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.