Skip to content
Conference

STGraphVQA: spatial-temporal graph reasoning with hierarchical cognition for interpretable driving scene understanding

Jul 2026 · International Conference on Image Processing and Intelligent Control · Vol 14262, pp. 142621I - 142621I-7 · 0 citations · 9 references
Engineering

TL;DR

STGraphVQA, a framework that represents driving scenes as dynamic spatiotemporal graphs, where nodes denote traffic participants, edges encode spatial and semantic relationships, and the temporal dimension captures their evolution, provides a promising direction toward interpretable autonomous driving systems.

Abstract

Visual question answering (VQA) in autonomous driving scenarios demands strong spatiotemporal reasoning capabilities, yet existing vision-language models lack explicit modeling of dynamic relationships in complex traffic scenes. We propose STGraphVQA, a framework that represents driving scenes as dynamic spatiotemporal graphs, where nodes denote traffic participants, edges encode spatial and semantic relationships, and the temporal dimension captures their evolution. A hierarchical reasoning architecture progressively processes information through perception, relation, and decision layers, simulating the human driving cognitive process. A logit-level constrained decoding mechanism further ensures that generated answers comply with traffic rules and physical feasibility. Experiments on DriveLM and STRIDE-QA demonstrate that STGraphVQA significantly outperforms state-of-the-art baselines, achieving a Top-1 accuracy of 76.8% and a reasoning chain completeness of 82.3%, providing a promising direction toward interpretable autonomous driving systems.

View source

Similar papers

Jul 2026

Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency

A cognitive dual-process planning framework that represents planning-relevant scene knowledge in a machine-parsable structured chain-of-thought (S-CoT) schema and shows how explicit scene knowledge can be operationalized through adaptive reasoning and rule-based verification to support high-level VLM planning decisions...

Zhongyao Yang, Haoyu Li, Yuchen Yan et al. · 0 citations
Preprint Aug 2026

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

Space Tokens is introduced, a lightweight, architecture-agnostic framework that equips VLMs with explicit continuous spatial representations without requiring additional inference-time modules, and demonstrates that continuous spatial tokens provide an effective, interpretable, and computationally efficient mechanism f...

Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling

MV-STRIDE, a Multi-View hierarchical SpaTial Reasoning dataset with Interdependent and DEcomposed capabilitiEs, explicitly models the dependency relationships between foundational perception, scene understanding, and complex contextual reasoning, providing a coherent learning pathway aligned with human spatial cognitio...

Jin Xu, Xiao-Jiang Huang, Zhuo Luo et al. · 0 citations
Preprint Aug 2026

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

This work designs a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs and innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations.

Jing Wu, Jianhua Wu, Jiayi Guan et al. · 0 citations
Preprint Aug 2026

ChronoVision: Temporal Reasoning via Latent State Reconstruction

This work proposes ChronoVision, a multimodal framework designed to align visual logic with latent imagery, and introduces Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task.

Yi-Fan Shen, Jian Xu, Boyi Li et al. · 1 citation
Preprint Aug 2026

GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model

GaussVLA is proposed, a Mamba-based VLA that incorporates two custom modules: Gaussian Spatial Tokenizer (GST) to lift frozen semantic and depth features into compact 3D Gaussian tokens, and Depth-Aware Chain-of-Thought (DA-CoT) that performs structured, non-autoregressive geometric reasoning under language and flow-ti...

MD SELIM SAROWAR, Md Tanvir Islam, Sungho Kim et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.