Skip to content

Visual Graph Reasoning via Knowledge Compilation

Sep 2026 · 0 citations
Computer Science

TL;DR

VGCompiler organizes reasoning around two compilers: a representation compiler that recovers a structure-preserving intermediate graph representation from visual input, and an operation compiler that compiles query intent under the recovered graph state into an executable graph operation.

Abstract

Visual graph reasoning requires answering graph-theoretic questions directly from graph images, where graph topology and state are conveyed visually rather than given in symbolic form. Despite recent progress of vision-language models (VLMs), current approaches to visual graph reasoning still fail on simple visual graph problems. This reveals a fundamental limitation of existing approaches: they prioritize final-answer supervision over the intermediate recovery of an explicit graph representation that preserves graph topology and state from visual input. To address this limitation, we propose VGCompiler, a compilation-centric paradigm for visual graph reasoning via knowledge compilation. VGCompiler organizes reasoning around two compilers: a representation compiler that recovers a structure-preserving intermediate graph representation from visual input, and an operation compiler that compiles query intent under the recovered graph state into an executable graph operation. Specifically, we build VGCompiler on Qwen3-VL-8B and train it with reinforcement learning guided by a layered reward over executability, compiled graph validity, representation quality, and operation quality. VGCompiler uses a frozen observer to summarize graph and question conditions into lightweight signatures, enabling archive retrieval and code reuse across similar regimes. Experiments on three benchmarks GVLQA, VisionGraph, and VGCURE, show that Qwen-VGCompiler, built on an 8B backbone, surpasses the strongest closed-source VLM baseline by 28.9% and the strongest code-based baseline by 23.7%, while maintaining high efficiency. We further evaluate VGCompiler on three real-world domains, including metro routing, logistics delivery, and network fault assessment, where it generalizes across heterogeneous visual graphs and domain-grounded tasks.

View source

Similar papers

#natural language process... Preprint Sep 2026

Question-Specific Knowledge Graphs for Efficient Visual Reasoning

Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs into textual representations. The effectiveness of this translation depends on how well visual details are retained; models need to surface and align both explicit and im...

Ting-Chih Chen, Emile van Krieken, Shu-Jian Yu et al. · 0 citations
#machine learning Conference Open access Sep 2026

When Vision Meets Graphs: A Survey on Graph Reasoning and Learning

This survey provides a first systematic overview of the emerging area of vision meets graphs, which treats visual depictions of graphs as first-class inputs for reasoning and learning, and organize existing work into three threads.

Xinjian Zhao, Wei Pang, Zhixuan Yu et al. · 2 citations
Preprint Aug 2026

TopoAgent: A Structure-Aware Perception-to-Reasoning Framework for Diagram-to-Graph Topology Extraction with Large Vision-Language Models

Diagram-to-graph topology extraction aims to extract a graph of entities and their connections from a structural diagram. This task remains challenging for current vision-language models because it requires both fine-grained perceptual grounding and topology-aware reasoning with global consistency. We present TopoBench...

Bang-Wei Guo, Xujiang Zhao, Yan-Chi Liu et al. · 0 citations
Preprint Sep 2026

A Topological Representation with Object-Path Graphs for Open-Vocabulary Instance Navigation

This work proposes an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation, and introduces a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly o...

Lin-Wei Zheng, Dao-Jie Peng, Bing-Tao Wang et al. · 0 citations
#small language model Preprint Sep 2026

Think-Verify-Revise: Neuro-Symbolic Visual Reasoning with Vision-Language Models and Dynamic Logic Tensor Networks

A Neuro-Symbolic framework that closes this gap by tightly coupling a Vision-Language Model for automatic First-Order Logic rule induction with a Dynamic Logic Tensor Network for differentiable rule verification, in a closed iterative feedback loop is proposed.

Homayoun Afshari, Pietro Basci, A. Russo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving

This work introduces a query-adaptive neuro-symbolic reasoning framework that explicitly allocates computation according to the nature of the query and shifts the role of the LLM from a universal reasoning engine to a targeted semantic reasoner, while allowing deterministic computation to be handled exactly and efficie...

Yi-Yao Wang, Pei Liu, Fang Liu et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.