A systematic survey and diagnostic evaluation of MLLMs for RSISU and demonstrates the strong transferability of general-purpose CV-MLLMs and shows that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks.
Abstract
The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.
UniRS-Instruct is presented, a high-quality, diversified, and unified multimodal instruction-following dataset for RSI understanding that unifies diverse tasks, including image captioning, visual question answering, visual grounding, and region-level captioning, into a consistent format.
Lin-Rui Xu, Yuhan Wang, Ling Zhao et al.· IEEE Journal of Selected Top...· 0 citations
This paper investigates whether meta-prompting with large language models (LLMs) can improve zero-shot scene classification in RS by automatically generating semantically rich class descriptions and highlights the potential of open-source LLMs as scalable prompt generators for zero-shot remote-sensing recognition.
Antonis Promponas, Eirini Baltzi, Valsamis Ntouskos et al.· The International Archives o...· 1 citation
Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it difficult to exploit heterogeneous sensors such as SAR, multi-spectral imagi...
Xian-Yang Miao, Ke-Lu Yao, Ye-Hua Huang et al.· 0 citations
GeoVP is proposed as a visual prompting MLLM for multisource RS image understanding, enabling unified image-level and region-level understanding under different prompt granularities and demonstrating the effectiveness of explicit visual prompting for prompt-conditioned region understanding in multisource RS imagery.
Le Yu, Yuan-Wen Wang, Xiao-Tong Qi· IEEE Journal of Selected Top...· 0 citations
RS-Florence is proposed, a compact unified model that addresses remote sensing perception systems through a Prompt-Driven Sequence-to-Sequence framework, which maps images and task-specific prompts into a unified sequence of natural language and discrete geometric tokens.
Yang Liu, Wei-Xing Luo, Huai-Zhou Qi et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.