Canvas Beyond Words: Teaching Machines to Sketch Empowers Their Spatial Intelligence.
Abstract
How to encode and evaluate spatial intelligence in foundation models remains an open challenge. Existing approaches often rely on textual proxies and VQA-style evaluation to assess visual-spatial intelligence (VSI), which can obscure geometric structure, encourage linguistic shortcuts, and hinder attribution to genuinely spatial reasoning abilities. To address this limitation, we introduce spatial intelligence grid (SIG), a structured grid-based schema that explicitly represents object layouts, inter-object relations, and physically grounded priors. As a complementary modality to text, SIG offers a faithful and compositional representation of scene structure for foundation-model reasoning. Building on SIG, we further propose SIG-informed evaluation metrics for measuring a model's intrinsic VSI, thereby disentangling spatial competence from language priors. In few-shot in-context learning (ICL) experiments with state-of-the-art multimodal large language models (MLLMs), including GPT- and Gemini-family models, SIG consistently delivers larger, more stable, and more comprehensive improvements across all VSI metrics than VQA-only representations. These results highlight its potential as an effective annotation and training schema for learning VSI. We further demonstrate the generalizability of SIG across diverse scenarios and multiple MLLMs, and systematically examine how data sample selection influences SIG-based ICL through extensive experiments. In addition, we release SIGBench, a benchmark comprising 1.4K driving frames annotated with ground-truth SIG labels and human gaze traces, supporting both grid-based machine VSI tasks and attention-driven, human-like VSI tasks in autonomous-driving settings.