Skip to content

MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling

Sep 2026 · 0 citations · 61 references
Computer Science

TL;DR

MV-STRIDE, a Multi-View hierarchical SpaTial Reasoning dataset with Interdependent and DEcomposed capabilitiEs, explicitly models the dependency relationships between foundational perception, scene understanding, and complex contextual reasoning, providing a coherent learning pathway aligned with human spatial cognition.

Abstract

Despite the rapid progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, robust multi-view spatial reasoning remains a fundamental bottleneck due to the lack of structured 3D cognitive pathways in existing datasets. To address this, we introduce MV-STRIDE, a Multi-View hierarchical SpaTial Reasoning dataset with Interdependent and DEcomposed capabilitiEs. Moving beyond flat data structures, MV-STRIDE explicitly models the dependency relationships between foundational perception, scene understanding, and complex contextual reasoning, providing a coherent learning pathway aligned with human spatial cognition. We develop a systematic QA generation pipeline leveraging diverse 3D scene sources that enforces cross-view dependency constraints to prevent single-view solvability, generating multi-level spatial reasoning tasks supported by cognitively grounded chain-of-thought supervision for complex inference. Extensive evaluations demonstrate that our multi-stage training framework based on our hierarchical dataset achieves state-of-the-art performance across multiple spatial reasoning benchmarks, notably the multi-view oriented MMSI-Bench. Our approach enables MLLMs to maintain robust, 3D-consistent spatial reasoning across diverse viewpoints. The code and dataset are available at https://co1dspring.github.io/MV-STRIDE/.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

Preliminary evidence is presented that reinforcement learning with verifiable rewards can elicit some latent multi-view competence in the base model, pointing to training-time approaches as a promising direction for future work.

Hyungjin Chung, Byeongjun Park, Joonseok Lee et al. · 0 citations
Preprint Aug 2026

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

It is shown that MVRD makes visual representations more geometric while retaining language alignment, and generalizes to 3D scene understanding tasks such as object grounding, dense captioning, and question answering, while approaching feature fusion methods with considerably fewer added parameters and lower latency.

K. T. Nguyen, Hanbo Shim, Jinwoo Kim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

This work introduces SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination and proposes a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulatio...

Soohyun Ryu, Sohee Kim, Eunho Yang · 0 citations
#computer vision Preprint Sep 2026

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close th...

Jaewoo Jung, Hyeonseo Yu, Honggyu An et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning

SpatialSpeak is introduced, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning and achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.

Yang Cao, Jia-Xin Zhang, Dave Zhenyu Chen et al. · 0 citations
Open access 2026

Decoupled global-local collaborative network for visual question answering

Visual Question Answering (VQA) aims to achieve cross-modal semantic understanding through joint modeling of visual content and natural language. Although existing attention-based approaches effectively align features, they struggle to simultaneously accommodate global semantic modeling and local fine-grained perceptio...

Gan-Long Zhou, Dezhi Han, Xiang Shen et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.