Preprint
Sep 2026
INTCORT: Training-Free Spatial Reasoning Enhancement for Vision-Language Models via Input Transformations and Confidence Routing
This work proposes INTCORT, a training-free spatial reasoning enhancement framework that constructs multiple inference views through input transformations and aggregates their predictions via relation-token confidence routing, without modifying the VLM's internal mechanisms.
Hao-Ran Sun, Jing-Qi Xu, Yan-Hui Li et al.
· 0 citations