Skip to content
Preprint

INTCORT: Training-Free Spatial Reasoning Enhancement for Vision-Language Models via Input Transformations and Confidence Routing

Sep 2026 · 0 citations · 21 references
Computer Science

TL;DR

This work proposes INTCORT, a training-free spatial reasoning enhancement framework that constructs multiple inference views through input transformations and aggregates their predictions via relation-token confidence routing, without modifying the VLM's internal mechanisms.

Abstract

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in multimodal tasks, yet they still exhibit poor ability in spatial reasoning. Existing training-dependent and training-free enhancement methods suffer from high computational costs with catastrophic forgetting and internal mechanism interference that compromises general capabilities, respectively. In this work, we first verify two key hypotheses: appropriate geometric image transformation and query-reversal transformation can recover incorrect spatial predictions, and correct predictions exhibit higher relation-token confidence than incorrect ones. Based on these findings, we propose INTCORT, a training-free spatial reasoning enhancement framework that constructs multiple inference views through input transformations and aggregates their predictions via relation-token confidence routing, without modifying the VLM's internal mechanisms. Experimental results on several commonly-used benchmarks demonstrate that INTCORT substantially improves spatial reasoning accuracy across diverse VLMs, achieving an average improvement of 10.01% over all models and benchmarks. Compared with prior works, INTCORT achieves superior performance with improvements of up to 25.01%.

View source

Similar papers

Preprint Aug 2026

GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model

GaussVLA is proposed, a Mamba-based VLA that incorporates two custom modules: Gaussian Spatial Tokenizer (GST) to lift frozen semantic and depth features into compact 3D Gaussian tokens, and Depth-Aware Chain-of-Thought (DA-CoT) that performs structured, non-autoregressive geometric reasoning under language and flow-ti...

MD SELIM SAROWAR, Md Tanvir Islam, Sungho Kim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Geometry-Aware Test-Time Learning for Quantitative Spatial Reasoning

Quantitative spatial reasoning in visual-language models (VLMs) aims to infer spatial distances and directional relationships among objects in 3D space from a 2D image and a natural language query. Despite recent progress, VLM spatial reasoning remains brittle under distribution shifts, largely due to the high cost of...

Ge-Ge Zhang, Shuai-Cheng Niu, Gang Dai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models

Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes...

Rafi Ibn Sultan, Xiang-Yu Zhou, Mohammad O. S. Chowdhury et al. · 0 citations
Preprint Aug 2026

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs

This work proposes a modular and training-free framework for spatial reasoning verification and correction, which associates atomic spatial evidence extracted from Chain-of-Thought reasoning with visual entities, spatial relations, source steps, and visual evidence.

Yang Yang, Jia-Wei Chen, Tai-Ran Chen et al. · 0 citations
Preprint Aug 2026

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

Space Tokens is introduced, a lightweight, architecture-agnostic framework that equips VLMs with explicit continuous spatial representations without requiring additional inference-time modules, and demonstrates that continuous spatial tokens provide an effective, interpretable, and computationally efficient mechanism f...

Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian et al. · 0 citations
Review Open access Aug 2026

Spatial intelligence in vision-language models: a comprehensive survey

This survey provides a comprehensive and unified overview of recent advances in spatial intelligence for VLMs, summarize core concepts behind spatial reasoning in VLMs, analyze why spatial failures occur, and organize existing solutions into a clear framework spanning prompting-based techniques, model improvements, exp...

Di-Sheng Liu, Tuo Liang, Zhe Hu et al. · 9 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.