Skip to content
Open access

UniRS-Instruct: A Principle-Guided Unified Instruction-Following Dataset for Remote Sensing Understanding

2026 · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · Vol 19, pp. 25589-25607 · 0 citations · 61 references

Abstract

Driven by multimodal large language models (MLLMs), remote sensing image (RSI) understanding is undergoing a paradigm shift, evolving from learning a domain-specific model to learning a general foundation model with domain adaptation (LaGD). Under the LaGD paradigm, conventional datasets, such as DOTA and RSICD, which fueled progress in RSI understanding over the past decade, are no longer adequate for emerging tasks because their annotation formats are task-specific and lack the language-level supervision required by MLLMs. We argue that a new dataset must be purposefully designed to support three core capabilities: First, generalization, enabling models to learn shared knowledge across tasks through a unified annotation format; Second, complex scene understanding, training models to capture fine-grained object attributes and spatial relationships and to describe scenes in detailed natural language; Finally, reasoning, equipping models with high-level visual reasoning through multiturn dialogues. To this end, we present UniRS-Instruct, a high-quality, diversified, and unified multimodal instruction-following dataset for RSI understanding. UniRS-Instruct unifies diverse tasks, including image captioning, visual question answering, visual grounding, and region-level captioning, into a consistent (question and answer) format. To construct fine-grained and context-aware instruction data, we propose a hierarchical prompting strategy: at the local level, objects are identified via rotated bounding boxes to describe their fine-grained attributes and spatial relationships; at the global level, local information is integrated with the full image to generate detailed scene-level instruction descriptions through GPT-4 V. Extensive experiments on multiple remote sensing benchmarks demonstrate that MLLMs fine-tuned with UniRS-Instruct achieve superior performance in image captioning, visual question answering, and visual grounding tasks, and exhibit stronger capabilities in describing fine-grained information, uncovering implicit knowledge, and conducting complex reasoning compared with models trained on existing datasets.

Read PDF