Vision-language models have achieved impressive progress, yet they still struggle with spatial intelligence–understanding where objects are, how they relate, and how space changes across viewpoints. This limitation matters for embodied AI, autonomous driving, and spatially consistent generation. Meanwhile, rapid advances in spatially enhanced VLMs have produced a scattered literature with inconsistent terminology, methods, and evaluation practices. In this survey, we provide a comprehensive and unified overview of recent advances in spatial intelligence for VLMs. We summarize core concepts behind spatial reasoning in VLMs, analyze why spatial failures occur, and organize existing solutions into a clear framework spanning prompting-based techniques, model improvements, explicit 2D cues, 3D enrichment, and data-driven strategies. We also examine how spatial ability is currently measured and report an empirical study across 37 models and 9 representative benchmarks. Our analysis highlights current best-performing approaches, clarifies when different strategies help or fail, shows the existence of performance gaps across different evaluation datasets and reveals the potential design biases in current spatial understanding benchmarks. By consolidating evidence and outlining open challenges, this survey offers a practical roadmap for building more spatially capable VLMs. We release our
evaluation code
and maintain a curated
paper repository
to support the rapidly growing research on spatial intelligence in vision-language models.
Disheng Liu, Tuo Liang, Zhe Hu et al.· Artificial Intelligence Revi...· 6 citations
Counterfactual Realignment (CoRe), a training-free framework that recovers a frozen VLA at inference time without failure data, is proposed, a training-free framework that recovers a frozen VLA at inference time without policy fine-tuning or failure-specific recovery training.
Yanyan Zhang, Disheng Liu, Kai Ye et al.· 0 citations