From Foundation Models to Geospatial Intelligence: A Taxonomy and Comparative Analysis for Smart City Remote Sensing
Abstract
Recent advances in foundation models have significantly reshaped remote sensing by enabling large-scale representation learning and improved transferability across a wide range of Earth observation tasks. However, previous reviews have largely focused on specific model families or applications, providing limited discussion of the capabilities and requirements of foundation models for smart-city scenarios. In this paper, we provide a comprehensive review of foundation models for remote sensing from a smart-city perspective. We first propose a taxonomy that classifies remote sensing foundation models into four groups: vision foundation models, vision-language models, segmentation foundation models, and geospatial multimodal models. We then compare these models in terms of learning approach, transferability, multimodal reasoning, computational cost, and applicability to smart-city scenarios, complemented by a qualitative evaluation of CLIP, BLIP-2, and LLaVA on aerial scenes. Finally, we summarize the open research issues, such as domain adaptation, efficient fine-tuning, multimodal fusion, explainability, and trustworthy AI, and suggest future directions toward autonomous, multimodal, and agentive geospatial intelligence systems for smart cities.