VGCompiler organizes reasoning around two compilers: a representation compiler that recovers a structure-preserving intermediate graph representation from visual input, and an operation compiler that compiles query intent under the recovered graph state into an executable graph operation.
Abstract
Visual graph reasoning requires answering graph-theoretic questions directly from graph images, where graph topology and state are conveyed visually rather than given in symbolic form. Despite recent progress of vision-language models (VLMs), current approaches to visual graph reasoning still fail on simple visual graph problems. This reveals a fundamental limitation of existing approaches: they prioritize final-answer supervision over the intermediate recovery of an explicit graph representation that preserves graph topology and state from visual input. To address this limitation, we propose VGCompiler, a compilation-centric paradigm for visual graph reasoning via knowledge compilation. VGCompiler organizes reasoning around two compilers: a representation compiler that recovers a structure-preserving intermediate graph representation from visual input, and an operation compiler that compiles query intent under the recovered graph state into an executable graph operation. Specifically, we build VGCompiler on Qwen3-VL-8B and train it with reinforcement learning guided by a layered reward over executability, compiled graph validity, representation quality, and operation quality. VGCompiler uses a frozen observer to summarize graph and question conditions into lightweight signatures, enabling archive retrieval and code reuse across similar regimes. Experiments on three benchmarks GVLQA, VisionGraph, and VGCURE, show that Qwen-VGCompiler, built on an 8B backbone, surpasses the strongest closed-source VLM baseline by 28.9% and the strongest code-based baseline by 23.7%, while maintaining high efficiency. We further evaluate VGCompiler on three real-world domains, including metro routing, logistics delivery, and network fault assessment, where it generalizes across heterogeneous visual graphs and domain-grounded tasks.
Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs into textual representations. The effectiveness of this translation depends on how well visual details are retained; models need to surface and align both explicit and im...
Ting-Chih Chen, Emile van Krieken, Shu-Jian Yu et al.· 0 citations
This survey provides a first systematic overview of the emerging area of vision meets graphs, which treats visual depictions of graphs as first-class inputs for reasoning and learning, and organize existing work into three threads.
Xinjian Zhao, Wei Pang, Zhixuan Yu et al.· Proceedings of the Thirty-Fi...· 2 citations
Diagram-to-graph topology extraction aims to extract a graph of entities and their connections from a structural diagram. This task remains challenging for current vision-language models because it requires both fine-grained perceptual grounding and topology-aware reasoning with global consistency. We present TopoBench...
Bang-Wei Guo, Xujiang Zhao, Yan-Chi Liu et al.· 0 citations
This work proposes an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation, and introduces a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly o...
Lin-Wei Zheng, Dao-Jie Peng, Bing-Tao Wang et al.· 0 citations
A Neuro-Symbolic framework that closes this gap by tightly coupling a Vision-Language Model for automatic First-Order Logic rule induction with a Dynamic Logic Tensor Network for differentiable rule verification, in a closed iterative feedback loop is proposed.
Homayoun Afshari, Pietro Basci, A. Russo et al.· 0 citations
This work introduces a query-adaptive neuro-symbolic reasoning framework that explicitly allocates computation according to the nature of the query and shifts the role of the LLM from a universal reasoning engine to a targeted semantic reasoner, while allowing deterministic computation to be handled exactly and efficie...
Yi-Yao Wang, Pei Liu, Fang Liu et al.· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.