Skip to content

From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving

Sep 2026 · 0 citations · 25 references
Computer Science

TL;DR

This work introduces a query-adaptive neuro-symbolic reasoning framework that explicitly allocates computation according to the nature of the query and shifts the role of the LLM from a universal reasoning engine to a targeted semantic reasoner, while allowing deterministic computation to be handled exactly and efficiently.

Abstract

Autonomous-driving question answering requires reasoning over structured scene information, yet existing vision-language approaches largely delegate heterogeneous reasoning operations to a single neural inference process. We argue that this uniform strategy overlooks a fundamental distinction: some queries admit exact symbolic solutions, while others require semantic interpretation. We introduce a query-adaptive neuro-symbolic reasoning framework that explicitly allocates computation according to the nature of the query. At its core is a hierarchical Spatiotemporal Scene Graph (STSG) that separates persistent object identities from frame-specific states and represents spatial relations and temporal transitions as explicit directed structures. Given a query, a symbolic executor first attempts to resolve it through exact graph operations; only when symbolic execution abstains is an LLM invoked for semantic reasoning. For these unresolved queries, query-conditioned graph retrieval and evidence filtering preserve relation direction, temporal locality, and object semantics, providing the LLM with compact and verified task-relevant evidence. This design shifts the role of the LLM from a universal reasoning engine to a targeted semantic reasoner, while allowing deterministic computation to be handled exactly and efficiently. We evaluate the framework on 5,916 NuScenes-QA questions across all ten scenes of nuScenes v1.0-mini under an oracle-perception setting. The complete system achieves 80.63 percent overall accuracy with GPT-5.4-mini, improving over the corresponding LLM-only configuration by 5.48 percentage points; with DeepSeek-V4-Flash, the improvement reaches 6.64 points. The largest gains occur on counting questions, with improvements of 10.20 and 12.61 points, respectively. These results show that selective reasoning improves both accuracy and inference efficiency.

View source

Similar papers

Open access Aug 2026

From Ambiguity to Execution: An Agentic Neuro-Symbolic Framework for Transforming Building Regulations into Deterministic Constraints

This research proposes an Agentic Neuro-Symbolic Framework that decouples semantic interpretation from geometric verification and establishes a scalable foundation for autonomous compliance, demonstrating that AI reliability in engineering significantly improves when probabilistic models orchestrate deterministic tools...

N. Mirhosseini, D. Shojaei, Soheil Sabri · 0 citations
Preprint Sep 2026

Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

A neuro-symbolic framework that combines learned VLA control with explicit task graphs and multimodal procedural memory is investigated, which positions structured symbolic reasoning and demonstration-derived visual guidance as complementary mechanisms for reliable long-horizon VLA manipulation.

Vivek Chavan, Ya-Huan Shi, O. Heimann et al. · 0 citations
Preprint Sep 2026

CASCADE: A Spatio-Temporal-Causal Reasoning Representation and Dataset for Driving

This work introduces CASCADE (Causal Spatio-Temporal Analysis of Driving Environments), a structured scene representation for reasoning in driving scenes and a human-annotated dataset built on it that provides the reference for benchmarking the reasoning abilities of Physical AI models, and verifying the quality of aut...

Jenny Schmalfuss, Despoina Paschalidou, Simon Gerstenecker et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning

Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonstrate that a pure Large Language Model (LLM), when...

Wei-Chen Dai, Rafael Cabral, Ziyi Shou et al. · 0 citations
Preprint Sep 2026

GraphPoint: Semantic Entity Graphs and Point Trajectories for Compositional Robot Manipulation

Robot manipulation policies often struggle to generalize beyond their demonstrations, even when new instructions involve familiar objects and behaviors. When language and scenes are strongly correlated during training, a policy can learn a fixed visual-action mapping rather than respond to the requested behavior. We in...

Kang-Ping Luo, He-Sheng Wang · 0 citations
Conference Aug 2026

Improving Visual Question Answering Via Rule- Guided Neuro-Symbolic Learning

Visual Question Answering (VQA) involves models combining reasoning through visual scenes and natural language questions and typically related to compositional and relational reasoning. Even though deep neural models have demonstrated high empirical results on VQA benchmarks, they are often based on implicit associatio...

Akash Badhan, Priyank Arora, Rishabh Garg et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.