Skip to content

Author

Disheng Liu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Aug 2026

Spatial intelligence in vision-language models: a comprehensive survey

Vision-language models have achieved impressive progress, yet they still struggle with spatial intelligence–understanding where objects are, how they relate, and how space changes across viewpoints. This limitation matters for embodied AI, autonomous driving, and spatially consistent generation. Meanwhile, rapid advances in spatially enhanced VLMs have produced a scattered literature with inconsistent terminology, methods, and evaluation practices. In this survey, we provide a comprehensive and unified overview of recent advances in spatial intelligence for VLMs. We summarize core concepts behind spatial reasoning in VLMs, analyze why spatial failures occur, and organize existing solutions into a clear framework spanning prompting-based techniques, model improvements, explicit 2D cues, 3D enrichment, and data-driven strategies. We also examine how spatial ability is currently measured and report an empirical study across 37 models and 9 representative benchmarks. Our analysis highlights current best-performing approaches, clarifies when different strategies help or fail, shows the existence of performance gaps across different evaluation datasets and reveals the potential design biases in current spatial understanding benchmarks. By consolidating evidence and outlining open challenges, this survey offers a practical roadmap for building more spatially capable VLMs. We release our evaluation code and maintain a curated paper repository to support the rapidly growing research on spatial intelligence in vision-language models.

Disheng Liu, Tuo Liang, Zhe Hu et al. · 6 citations
Preprint Aug 2026

Imagining Recovery: Inference-Time Counterfactual Realignment for Vision-Language-Action Models

Counterfactual Realignment (CoRe), a training-free framework that recovers a frozen VLA at inference time without failure data, is proposed, a training-free framework that recovers a frozen VLA at inference time without policy fine-tuning or failure-specific recovery training.

Yanyan Zhang, Disheng Liu, Kai Ye et al. · 0 citations