Preprint
Aug 2026
Investigating Relational Reasoning in VLMs
This work uses the Qwen3-VL-4B (Bai et al., 2025), a modern VLM, to decode how visual information is encoded across depths and proposes a synthetic dataset of simple geometric shapes for controlled analysis, along with queries crafted to precisely test language cues.
Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan et al.
· 0 citations