Skip to content
Review

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

Jul 2026 · arXiv.org · Vol abs/2607.20284 · 0 citations · 87 references
Computer Science

TL;DR

A systematic survey and diagnostic evaluation of MLLMs for RSISU and demonstrates the strong transferability of general-purpose CV-MLLMs and shows that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks.

Abstract

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.

View source

Similar papers

Open access 2026

UniRS-Instruct: A Principle-Guided Unified Instruction-Following Dataset for Remote Sensing Understanding

UniRS-Instruct is presented, a high-quality, diversified, and unified multimodal instruction-following dataset for RSI understanding that unifies diverse tasks, including image captioning, visual question answering, visual grounding, and region-level captioning, into a consistent format.

Lin-Rui Xu, Yuhan Wang, Ling Zhao et al. · 0 citations
Open access Jul 2026

Meta-Prompting with Open-Source Language Models for Zero-Shot Scene Classification in Remote Sensing

This paper investigates whether meta-prompting with large language models (LLMs) can improve zero-shot scene classification in RS by automatically generating semantically rich class descriptions and highlights the potential of open-source LLMs as scalable prompt generators for zero-shot remote-sensing recognition.

Antonis Promponas, Eirini Baltzi, Valsamis Ntouskos et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data

Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it difficult to exploit heterogeneous sensors such as SAR, multi-spectral imagi...

Xian-Yang Miao, Ke-Lu Yao, Ye-Hua Huang et al. · 0 citations
Open access 2026

GeoVP: A Unified Visual Prompting Framework for Multisource Remote-Sensing Image Understanding

GeoVP is proposed as a visual prompting MLLM for multisource RS image understanding, enabling unified image-level and region-level understanding under different prompt granularities and demonstrating the effectiveness of explicit visual prompting for prompt-conditioned region understanding in multisource RS imagery.

Le Yu, Yuan-Wen Wang, Xiao-Tong Qi · 0 citations
Conference Open access Sep 2026

Unified Sequence Modeling for Remote Sensing: A Parameter-Efficient Foundation Model via Prompt-Driven Granularity Alignment

RS-Florence is proposed, a compact unified model that addresses remote sensing perception systems through a Prompt-Driven Sequence-to-Sequence framework, which maps images and task-specific prompts into a unified sequence of natural language and discrete geometric tokens.

Yang Liu, Wei-Xing Luo, Huai-Zhou Qi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.