This work introduces a query-adaptive neuro-symbolic reasoning framework that explicitly allocates computation according to the nature of the query and shifts the role of the LLM from a universal reasoning engine to a targeted semantic reasoner, while allowing deterministic computation to be handled exactly and efficiently.
Abstract
Autonomous-driving question answering requires reasoning over structured scene information, yet existing vision-language approaches largely delegate heterogeneous reasoning operations to a single neural inference process. We argue that this uniform strategy overlooks a fundamental distinction: some queries admit exact symbolic solutions, while others require semantic interpretation. We introduce a query-adaptive neuro-symbolic reasoning framework that explicitly allocates computation according to the nature of the query. At its core is a hierarchical Spatiotemporal Scene Graph (STSG) that separates persistent object identities from frame-specific states and represents spatial relations and temporal transitions as explicit directed structures. Given a query, a symbolic executor first attempts to resolve it through exact graph operations; only when symbolic execution abstains is an LLM invoked for semantic reasoning. For these unresolved queries, query-conditioned graph retrieval and evidence filtering preserve relation direction, temporal locality, and object semantics, providing the LLM with compact and verified task-relevant evidence. This design shifts the role of the LLM from a universal reasoning engine to a targeted semantic reasoner, while allowing deterministic computation to be handled exactly and efficiently. We evaluate the framework on 5,916 NuScenes-QA questions across all ten scenes of nuScenes v1.0-mini under an oracle-perception setting. The complete system achieves 80.63 percent overall accuracy with GPT-5.4-mini, improving over the corresponding LLM-only configuration by 5.48 percentage points; with DeepSeek-V4-Flash, the improvement reaches 6.64 points. The largest gains occur on counting questions, with improvements of 10.20 and 12.61 points, respectively. These results show that selective reasoning improves both accuracy and inference efficiency.
This research proposes an Agentic Neuro-Symbolic Framework that decouples semantic interpretation from geometric verification and establishes a scalable foundation for autonomous compliance, demonstrating that AI reliability in engineering significantly improves when probabilistic models orchestrate deterministic tools...
N. Mirhosseini, D. Shojaei, Soheil Sabri· Buildings· 0 citations
A neuro-symbolic framework that combines learned VLA control with explicit task graphs and multimodal procedural memory is investigated, which positions structured symbolic reasoning and demonstration-derived visual guidance as complementary mechanisms for reliable long-horizon VLA manipulation.
Vivek Chavan, Ya-Huan Shi, O. Heimann et al.· 0 citations
This work introduces CASCADE (Causal Spatio-Temporal Analysis of Driving Environments), a structured scene representation for reasoning in driving scenes and a human-annotated dataset built on it that provides the reference for benchmarking the reasoning abilities of Physical AI models, and verifying the quality of aut...
Jenny Schmalfuss, Despoina Paschalidou, Simon Gerstenecker et al.· 0 citations
Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonstrate that a pure Large Language Model (LLM), when...
Wei-Chen Dai, Rafael Cabral, Ziyi Shou et al.· 0 citations
Robot manipulation policies often struggle to generalize beyond their demonstrations, even when new instructions involve familiar objects and behaviors. When language and scenes are strongly correlated during training, a policy can learn a fixed visual-action mapping rather than respond to the requested behavior. We in...
Visual Question Answering (VQA) involves models combining reasoning through visual scenes and natural language questions and typically related to compositional and relational reasoning. Even though deep neural models have demonstrated high empirical results on VQA benchmarks, they are often based on implicit associatio...
Akash Badhan, Priyank Arora, Rishabh Garg et al.· International Conference on...· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.