Skip to content

CRISP: Cross-Modal Residual Guidance and Spatial Realignment for Remote Sensing Visual Question Answering

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5636413-5636413 · 0 citations · 71 references

Abstract

Remote sensing visual question answering (RSVQA) remains challenging because the evidence relevant to the question in remote sensing (RS) imagery is typically sparse, spatially dispersed, and highly variable in scale and layout. This structural heterogeneity poses a major challenge to existing transfer strategies, which often struggle to achieve effective query-conditioned semantic focusing and localized spatial refinement. To address this issue, we propose CRISP, a task-specific parameter-efficient adaptation framework for RSVQA built on a frozen ViLT backbone. CRISP comprises two complementary components. First, a cross-modal residual guidance (CMRG) module generates instance-specific guidance tokens from pooled image and question summaries, steering early cross-modal interaction toward query-relevant content while suppressing background interference. Second, an attention-guided spatial realignment (ASR) module performs offset-guided feature realignment within intermediate Transformer layers, enabling localized refinement of spatial evidence under scale variation and sparse semantic distribution. Extensive experiments on the RSVQA-LR and RSVQA-HR benchmarks show that CRISP achieves strong overall performance and consistently improves overall accuracy (OA) and average accuracy (AA) over prior methods, with particularly notable gains on presence, comparison, and region-related questions. These results demonstrate that residual guidance and spatial realignment together provide an effective task-specific parameter-efficient adaptation strategy for RSVQA under the frozen-backbone setting. The code will be available at https://github.com/PhD-Xu/CRISP

View source

Similar papers

Conference 2026

Cross-Modal Dynamic Aggregation with Adaptive Relevance Modulation Fusion Network for Remote Sensing Visual Question Answering

Remote Sensing Visual Question Answering (RS VQA) task aims to provide accurate answers to questions about RS images. However, the semantic gap between low-level visual features and high-level semantics complicates the understanding of complex questions. Moreover, the lack of dynamic modulation mechanisms for integrati...

Zi-Hua Zuo · 0 citations
Preprint Sep 2026

VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...

Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al. · 0 citations
Preprint Sep 2026

Selective Tool Use for Agentic Change Visual Question Answering in Remote Sensing

A selective tool use framework in which a single VLM either answers directly or invokes a deterministic change analysis tool to obtain question specific evidence to demonstrate the benefit of question-specific semantic evidence for Change VQA, while highlighting the influence of semantic prediction quality on the resul...

Y. Bazi, M. M. Al Rahhal, M. Mekhtiche et al. · 0 citations
#computer vision Preprint Sep 2026

From Coordinates to Candidate Regions: Temporal Change Localization via Region Selection in Remote Sensing Multimodal LLMs

Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual question answering over satellite imagery, yet localizing specific objects or changed regions remains challenging. Existing approaches rely on generating bounding box coordinates as token sequences, which is fragile...

J. Chung, Sungjune Park, Yeongyun Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.