Skip to content

When Vision Meets Graphs: A Survey on Graph Reasoning and Learning

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · 2 citations · 63 references
Computer Science

TL;DR

This survey provides a first systematic overview of the emerging area of vision meets graphs, which treats visual depictions of graphs as first-class inputs for reasoning and learning, and organize existing work into three threads.

Abstract

Graphs are a fundamental data structure underlying many problems in the natural and social sciences. Over the past decade, Graph Neural Networks (GNNs) have dominated graph machine learning, supported by solid theoretical foundations. Yet scientists often understand graph structure through vision: chemists read molecular diagrams and social scientists inspect network visualizations. Despite decades of work on graph visualization, most graph learning pipelines still treat graphs purely as symbolic structures, rarely leveraging the visual form of graphs. We argue that this gap deserves renewed attention in the era of powerful vision and vision language models. This survey provides a first systematic overview of the emerging area we term vision meets graphs, which treats visual depictions of graphs as first-class inputs for reasoning and learning. We organize existing work into three threads. Vision for Graph Reasoning studies how models can use visual depictions of graphs to understand structure and carry out multi-step reasoning. Vision for Graph Learning explores how visual features can complement or augment graph encoders beyond known limitations of message passing. Scientific Graphs examines domains where standardized depiction conventions support both reasoning and learning. Our goal is to clarify what current methods can and cannot do, and to outline a path toward foundation models that perceive and reason about graphs as scientists do.

Read PDF

Similar papers

Preprint Aug 2026

GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models

GraphVerse is introduced, a unified benchmark that jointly evaluates perception, visual reasoning, and text-based graph reasoning in MLLMs under both single-image and paired-image settings and proposes VGR-Score, a process-sensitive metric that evaluates reasoning quality beyond final-answer accuracy.

Yuan-Fu Sun, Yuanhang Ren, Kang Li et al. · 1 citation · ⚡1
#artificial intelligence Preprint Sep 2026

Visual Graph Reasoning via Knowledge Compilation

VGCompiler organizes reasoning around two compilers: a representation compiler that recovers a structure-preserving intermediate graph representation from visual input, and an operation compiler that compiles query intent under the recovered graph state into an executable graph operation.

Rong-Zheng Wang, Zhe Wang, Ke Qin et al. · 0 citations
#graph neural networks Review Open access Sep 2026

Automated graph construction and graph neural network search: a survey

Graphs are widely used to describe objects and their interactions in physically-informed real-world networking scenario including transportation, networking and energy, etc. Graph neural network (GNN) is the latest deep learning (DL) model for processing graph-structured data, widely applied in various tasks, e.g., p...

Yu-Feng Wang, Xin-Ying-Jian-Gan-Zhi-De-Shen-Jing-Jia-Gou-Sou-Suo Wang, Jian-Hua Ma et al. · 0 citations
#machine learning Review Sep 2026

From topology learning to graph generation: A unifying perspective

Learning graph structures from data is a fundamental problem that spans a wide range of signal processing and machine learning tasks. While significant effort has been made to tackle the problem, existing research has largely evolved along two parallel directions. The first seeks to infer the topology of an individual...

Xiaowen Dong, Hoi-To Wai, Si-Heng Chen et al. · 0 citations
Preprint Aug 2026

ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion

ViSR-KGC, a visual subgraph reasoning approach for KGC, integrates three complementary capabilities to capture semantic correlations: identifying global topology dependencies via representation learning, analyzing local multimodal evidence using VLMs, and providing necessary commonsense knowledge inherent in pre-traine...

Jiafan Li, Meng-Xue Yang, Jiaqi Zhu et al. · 0 citations

Related blog posts

Microsoft Research Blog Aug 11, 2026

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.