Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· 0 citations· 37 references
TL;DR
RS-Florence is proposed, a compact unified model that addresses remote sensing perception systems through a Prompt-Driven Sequence-to-Sequence framework, which maps images and task-specific prompts into a unified sequence of natural language and discrete geometric tokens.
Abstract
Current remote sensing (RS) perception systems suffer from task heterogeneity, necessitating distinct architectures for classification, localization, and reasoning. While vision--language models (VLMs) offer a route toward unification, their computational cost can hinder deployment. In this work, we propose RS-Florence, a compact unified model that addresses these tasks through a Prompt-Driven Sequence-to-Sequence framework. Unlike traditional approaches that segregate semantic understanding and geometric localization, RS-Florence maps images and task-specific prompts into a unified sequence of natural language and discrete geometric tokens. This formulation helps bridge high-level semantics and low-level pixel perception. Experiments across 4 task families and 8 benchmarks show that our 0.23B model remains competitive with task-specific specialists. These results also suggest that multitask joint training improves performance on several benchmarks over single-task fine-tuning, indicating that a shared prompt-driven interface can serve both language and geometry tasks.
UniRS-Instruct is presented, a high-quality, diversified, and unified multimodal instruction-following dataset for RSI understanding that unifies diverse tasks, including image captioning, visual question answering, visual grounding, and region-level captioning, into a consistent format.
Lin-Rui Xu, Yuhan Wang, Ling Zhao et al.· IEEE Journal of Selected Top...· 0 citations
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...
Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al.· 0 citations
GeoVP is proposed as a visual prompting MLLM for multisource RS image understanding, enabling unified image-level and region-level understanding under different prompt granularities and demonstrating the effectiveness of explicit visual prompting for prompt-conditioned region understanding in multisource RS imagery.
Le Yu, Yuan-Wen Wang, Xiao-Tong Qi· IEEE Journal of Selected Top...· 0 citations
Prompt-driven vision-language models (VLMs) hold immense promise for accelerating dense remote sensing (RS) annotation, but static models suffer from severe performance degradation when deployed on novel scenes, unseen categories, or visually confusing backgrounds. Moreover, existing unified paradigms primarily rely on...
Kunquan Zhang, Peilang Li, Xikun Hu et al.· 0 citations
Multimodal vision–language modeling has emerged as a promising paradigm for remote sensing (RS) image understanding. However, existing methods are limited by fixed-resolution visual processing and single-scale vision–language alignment, making it difficult to simultaneously preserve fine-grained details and maintain se...
Siyu Zhang, Ying Chen, Lianlei Shan et al.· IEEE Transactions on Geoscie...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.