Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?
This work investigates two representative model families, LLaVA-1.5 and Qwen2.5, and provides a token-, layer-, and head-level account of how VLMs transform object grounding into spatial relations, showing that knowing where objects are is not equivalent to knowing how they relate.
Xiwei Liu, Yulong Li, Xinlin Zhuang et al.
· 0 citations