BSR-Bench: Benchmarking MLLMs’ Basic Spatial Reasoning Capabilities
Abstract
In this paper, we present BSR-Bench, a 2D and 3D spatial reasoning benchmark designed to probe the extent to which contemporary MLLMs can support core spatial cognitive processes. In the first part of the research, we examine whether AI models can support core spatial cognitive processes and creative interaction with the physical world, using origami as a test case for spatial cognition. We then create BSR-Bench which evaluates the spatial reasoning capabilities of three widely used commercial-level MLLMs across four fundamental domains: navigation, composition, relationships, and transformation. These four areas encompass 13 prompts and 18 benchmarking metrics addressing spatial reasoning tasks, including block decomposition, orientation, and transformation, and are evaluated based on (partial or full) accuracy and consistency. These findings suggest that, despite recent advances, current MLLMs remain more adept at logical or linguistic abstraction than at the spatially grounded reasoning that supports human cognition and creative interaction with the physical world.