2026· IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing· Vol 19, pp. 26739-26757· 0 citations· 109 references
TL;DR
GeoVP is proposed as a visual prompting MLLM for multisource RS image understanding, enabling unified image-level and region-level understanding under different prompt granularities and demonstrating the effectiveness of explicit visual prompting for prompt-conditioned region understanding in multisource RS imagery.
Abstract
Recent advances in prompt learning and multimodal large language models (MLLMs) have improved interactive image understanding. However, fine-grained remote-sensing (RS) interpretation remains challenging because text-only instructions are often insufficient to precisely specify regions of interest in complex scenes, and visual prompting methods developed for natural images generalize poorly to heterogeneous RS data. To address these challenges, GeoVP is proposed as a visual prompting MLLM for multisource RS image understanding. GeoVP supports point, box, and free-form prompts, enabling unified image-level and region-level understanding under different prompt granularities. It employs a hybrid vision encoder to extract multiscale semantic and structural features and a region-aware encoder to convert heterogeneous prompts into unified region representations. These cues are integrated with language instructions for fine-grained RS reasoning. A one-stage training strategy is adopted to improve cross-domain adaptation across natural-image and RS domains. In addition, an auxiliary Pixel-Level Localization Module provides qualitative mask-based visualization cues for prompted regions. GeoVP-650 K is constructed as a 654K-scale image–prompt–text triplet dataset covering optical, synthetic aperture radar, and infrared imagery. GeoVP achieves an average zero-shot classification accuracy of 79.05% on AID and an average cross-task classification accuracy of 87.55% on UCMerced. It also obtains an average semantic intersection over union of 98.25% on DIOR-RSVG under box prompts, demonstrating the effectiveness of explicit visual prompting for prompt-conditioned region understanding in multisource RS imagery.
Experiments on the RSICD and NWPU-Captions datasets demonstrate that DSRAT achieves state-of-the-art performance across six metrics on RSICD and all seven metrics on NWPU-Captions, validating the effectiveness of the proposed approach.
Rong Chen, Guo-Rui Ma, Lunjun Fan· The International Archives o...· 0 citations
Recent advances in remote sensing (RS) vision-language foundation models (VLFMs) rely heavily on large-scale paired image–text data. However, existing dataset construction methods mainly emphasize data scale while overlooking a more fundamental limitation: the lack of structured and comprehensive semantic representatio...
Yi-Guo He, Jun-Jie Zhu, Jun Wang et al.· IEEE Transactions on Geoscie...· 0 citations
RS-Florence is proposed, a compact unified model that addresses remote sensing perception systems through a Prompt-Driven Sequence-to-Sequence framework, which maps images and task-specific prompts into a unified sequence of natural language and discrete geometric tokens.
Yang Liu, Wei-Xing Luo, Huai-Zhou Qi et al.· Proceedings of the Thirty-Fi...· 0 citations
The emergence of large-scale vision–language models (VLMs) has significantly advanced remote sensing image–text retrieval (RSITR) by providing powerful cross-modal semantic priors. However, when adapted to the remote sensing (RS) domain, these models struggle to capture fine-grained representations due to their inheren...
Wen-Liang Du, Xiao-Yu Xu, Jia-Qi Zhao et al.· IEEE Transactions on Geoscie...· 0 citations
UniRS-Instruct is presented, a high-quality, diversified, and unified multimodal instruction-following dataset for RSI understanding that unifies diverse tasks, including image captioning, visual question answering, visual grounding, and region-level captioning, into a consistent format.
Lin-Rui Xu, Yuhan Wang, Ling Zhao et al.· IEEE Journal of Selected Top...· 0 citations
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...
Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.