Foundation Models for Remote Sensing Semantic Segmentation: A Review of Architectures, Adaptations, and Prospects
Abstract
Highlights What are the main findings? Review of four emerging foundation model paradigms for remote sensing image segmentation—Transformer-based architectures, state space models (Mamba), prompt-driven segmentation (SAM), and self-supervised or multimodal pre-training analyzing trade-offs in global context modeling, computational efficiency, and cross-modal representation. Synthesis of downstream adaptation strategies, including parameter-efficient fine-tuning (LoRA, adapters), prompt engineering, few-shot and zero-shot learning, open-vocabulary segmentation, and domain adaptation, revealing how each strategy addresses the gap between pre-training and remote sensing requirements. What are the implications of the main findings? Identification of fundamental bottlenecks limiting current models, including the tension between representation generality and remote sensing-specific adaptation, multimodal sensor heterogeneity, and insufficiencies in existing evaluation ecosystems and annotation paradigms. A forward-looking research roadmap toward remote-sensing-native pre-training, lightweight edge-deployable architectures, and unified open-world geospatial foundation models, providing guidance for future research and practical deployment. Abstract Remote sensing image segmentation is a foundational task in Earth observation. With the rapid growth of remote sensing datasets in terms of scale, modality diversity, semantic openness, and spatio-temporal complexity, the field is evolving from task-specific supervised learning toward foundation-model paradigms. Recent advances in foundation models—including Transformer-based architectures, Mamba-based state space models (SSMs), prompt-driven frameworks such as the Segment Anything Model (SAM), and self-supervised or multimodal pre-training—have profoundly reshaped the technical landscape of remote sensing image segmentation. This paper reviews recent progress from the perspectives of dataset evolution, model architectures, and downstream adaptation strategies, covering parameter-efficient fine-tuning, prompt engineering, few-shot and zero-shot learning, open-vocabulary segmentation, and domain adaptation. We further analyze core challenges including the tension between representation generality and remote sensing-specific adaptation, multimodal sensor heterogeneity, and the insufficiency of existing evaluation ecosystems. Finally, we discuss future directions toward remote-sensing-native pre-training, lightweight edge deployment, and unified open-world geospatial foundation models.