Skip to content

LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

LADDER is a novel framework that bridges diffusion language modeling with GraphRAG through graph-guided parallel decoding through graph-guided parallel decoding, and proposes an event-driven self-clocking retrieval, inspired by the key insight that 88% of target entities emerge early in the partially denoised state.

Abstract

Graph Retrieval-Augmented Generation (GraphRAG) has remarkably enhanced large language models on complex reasoning by leveraging structured entity topologies. However, existing frameworks heavily rely on standard autoregressive language models where the nature of inherent sequential generation severely hinders overall inference efficiency. Inspired by Diffusion Language Models (DLMs) that offer massive parallelism via continuous refine-in-parallel decoding, we aim to accelerate GraphRAG in the discrete space. However, it remains non-trivial for two challenges. First, partially denoised drafts are highly dynamic and uncertain, making dynamic graph grounding non-trivial. Second, raw denoising states are inherently noisy and unstable, making synchronous graph retrieval and multi-hop aggregation computationally prohibitive. To this end, we present LADDER, a novel framework that bridges diffusion language modeling with GraphRAG through graph-guided parallel decoding. Specifically, (i) we propose an event-driven self-clocking retrieval, inspired by our key insight that 88% of target entities emerge early in the partially denoised state, leading final commitment by an average of 5.7-9.6 steps. This mechanism dynamically triggers graph retrieval only when the set of graph-linkable entities expands, yielding an asynchronous self-clocking policy that bypasses learned gates or heuristic thresholds. (ii) An incomplete-query graph propagation module is designed to process the newly emerging entity queries using a specialized graph foundation model, continuously aggregating multi-hop evidence to sharpen parallel predictions and accelerate overall decoding convergence. Extensive experiments on three challenging multi-hop QA benchmarks show that LADDER raises average exact match from 39.6% to 45.2% while achieving a 4.1x latency reduction.

View source

Similar papers

Open access Sep 2026

From Topology to Cognition: A Unified Graph-Driven Framework for Scalable Intelligent Data Mining

Modern data are increasingly represented and utilized as interconnected networks, including collaboration graphs, multi-relational user–item interactions, and schema-less knowledge graphs that support retrieval-augmented generation pipelines. While graph analytics has become a key foundation for intelligent data mining...

Yao Hu, Qian Huang · 0 citations
Review Open access Sep 2026

Hybrid Graph Retrieval-Augmented Language Agents for Collaborative Recommendation

Recent advances in large language model (LLM) agents have shown promise for autonomous decision-making in recommender systems. However, existing approaches suffer from two fundamental limitations: flat agent memories that conflate different information modalities and prohibitive computational costs that prevent scaling...

Ivan Bulychev, Andrey V. Savchenko · 0 citations
#artificial intelligence Preprint Sep 2026

GraMRAG: Orchestrating Multi-Agent Multi-Step Reasoning via Graph Memory with Reinforcement Learning

Although existing multi-agent Retrieval-Augmented Generation (RAG) systems have demonstrated promise on complex multimodal reasoning tasks, they remain fundamentally limited in reasoning depth and memory structure, suffering from inadequate retrieval and state blindness when answering knowledge-intensive questions. To...

Zhong-Yu Wang · 0 citations
Preprint Aug 2026

Noesis: Bidirectional Graph-RAG with Adaptive Parallelism and Cross-Knowledge-Base Semantic Discovery

Noesis, a decoupled Graph-RAG architecture addressing limitations through four algorithms: Bidirectional Graph Traversal with a Graph-Feedback Context Resolver simulating human reading with degrading memory, an AIMD Concurrency Controller adapted from TCP congestion control, and Moesis, domain-aware selective quantizat...

Nicola Cogotti · 1 citation
#natural language process... Preprint Sep 2026

DA-DLM: Explicitly Modeling Token Dependencies in Diffusion Language Models

Diffusion Language Models (DLMs) generate text by iteratively denoising a masked sequence, independently predicting multiple tokens at each step. This conditional independence discards inter-token dependencies and degrades coherence-an issue that parallels the multi-modality problem in Non-Autoregressive Translation (N...

Peng-Yu Ji, Zi-Chen Zhang, Xiang Hu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.