Skip to content
Review Open access

Geospatial Vision-Language Models for Spatial Reasoning and Temporal Change Understanding: A Task-Centered Benchmarking Framework and Evidence Synthesis Discussion

Aug 2026 · International Journal of Information Technology and Computer Science Applications · 0 citations · 32 references

TL;DR

A rigorous review paper anchored in public benchmark evidence and a reference architecture for trustworthy geospatial VLMs that is centered on multi-temporal supervision, interactive change analysis, uncertainty-aware outputs, and evaluation protocols that measure not only accuracy but also transfer, calibration, and operational feasibility.

Abstract

Geospatial vision-language models (VLMs) are increasingly expected to do more than assign scene labels or generate generic captions. In realistic Earth-observation and urban intelligence settings, useful multimodal systems must support fine-grained spatial reasoning, cross-view interpretation, grounded localization, and explicit understanding of change across time. Yet the current literature remains fragmented across remote sensing visual question answering, visual grounding, urban multi-view reasoning, and bi-temporal change captioning. As a result, claims about progress are often task-local, benchmark-specific, and difficult to compare. This paper reconstructs the field around a more defensible technical center: geospatial multimodal intelligence as the joint problem of spatial reasoning and temporal change understanding. Rather than presenting unverifiable new benchmark runs, we develop a rigorous review paper anchored in public benchmark evidence and a reference architecture for reproducible future work. We first formalize a task-centered problem definition that unifies image-level, region-level, cross-view, and bi-temporal reasoning. We then propose a reference GST-VLM architecture consisting of spatial encoding, temporal difference modeling, multimodal fusion, task-specific decoding, and reliability estimation. Next, we synthesize publicly reported evidence from representative datasets and benchmarks including RSVQA, EarthVQA, VRSBench, GeoChat, LEVIR-CD, LEVIR-CC, SECOND-CC, CHOICE, GEOBench-VLM, CityBench, and UrBench. The synthesis shows that recent models are improving rapidly but remain far from robust geospatial reasoning systems: on GEOBench-VLM, the best public model reported only 41.7% multiple-choice accuracy; on UrBench, even GPT-4o still trails human performance by an average 17.4 percentage points; and while specialized systems such as GeoReasoner, GeoChat, GeoLLaVA, and MModalCC outperform generic baselines on targeted tasks, their gains remain strongly benchmark-dependent. Based on this evidence, we identify the principal bottlenecks as benchmark fragmentation, weak temporal grounding, inadequate calibration, scarce cross-region validation, limited deployment reporting, and insufficient integration of geometry with language-conditioned reasoning. The paper concludes with a concrete research agenda for trustworthy geospatial VLMs that is centered on multi-temporal supervision, interactive change analysis, uncertainty-aware outputs, and evaluation protocols that measure not only accuracy but also transfer, calibration, and operational feasibility.

Read PDF

Similar papers

Preprint Aug 2026

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

This work introduces SPATIALQUERY, a training- free framework for CIDQ reasoning from a single RGB image, together with SPATIALQUERY-1M, a benchmark containing over one million RGB-only question-answer pairs from 200 indoor scenes, and proposes Uncertainty-Aware Chain-of-Thought (UA-CoT) prompting, which incorporates g...

Hai-Tra Nguyen, Tung Vu, Cong Tran · 0 citations
Preprint Aug 2026

CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment

Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the...

Yutian Jiang, Jiabo Liu, Xixuan Hao et al. · 1 citation
Preprint Aug 2026

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

This work introduces GeoMEB, a large-scale multimodal embedding benchmark that standardizes 45 urban evaluation tasks across retrieval, visual question answering, change detection, classification, and visual grounding, and presents Geo-Embed, a unified embedding model that adapts a shared vision-language backbone to in...

Jia-Peng Li, Yong Li, Jun-Jie Zhou et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Perceive to Hypothesize, Verify to Ground: An Agentic Reasoning Framework for Open-World Geo-Localization

This work reformulate geo-localization as a human-like perceive-then-verify reasoning problem and proposes GeoPAVE (Geo-localization Perception-and-Verification-Engine), a bi-level agentic framework that contains perception-based hypothesis generation via single-pass rollouts and verification-based evidence grounding f...

Yutian Jiang, Rui-Ji Li, Sisuo Lyu et al. · 0 citations
Preprint Aug 2026

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short seque...

Yupan Ding, Jing Xiao, Zhenyuan Zhang et al. · 0 citations
#computer vision Preprint Aug 2026

GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation

This work introduces GeoAgent, an agentic environment-based benchmark that requires agents to navigate Street View environments to refine their geolocalization through sequential reasoning, and establishes the challenges of embodied navigation and geospatial reasoning.

Arka Mukherjee, Soham Roy, Kartikeya Trivedi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.