Canvas Beyond Words: Teaching Machines to Sketch Empowers Their Spatial Intelligence.
How to encode and evaluate spatial intelligence in foundation models remains an open challenge. Existing approaches often rely on textual proxies and VQA-style evaluation to assess visual-spatial intelligence (VSI), which can obscure geometric structure, encourage linguistic shortcuts, and hinder attribution to genuine...