Jul 2026· International Conference on Generative Artificial Intelligence and Image Processing· Vol 14292, pp. 142920E - 142920E-6· 0 citations· 15 references
Engineering
TL;DR
DLPP (Diffusion-based Learning Path Planning Framework), which recasts path generation as a two-stage process of instructional planning followed by semantic instantiation, outperforms RL-Path and DiffPath baselines while remaining competitive with rule-based methods.
Abstract
Generating personalized learning paths remains an open problem in intelligent education systems, where conventional sequence recommendation methods typically output ranked resource ID lists without encoding pedagogical intent or supporting post-deployment refinement. To address these limitations, we present DLPP (Diffusion-based Learning Path Planning Framework), which recasts path generation as a two-stage process of instructional planning followed by semantic instantiation. In the first stage, a conditional diffusion model produces structured activity-type sequences within a continuous embedding space; nearest-neighbor quantization then maps these sequences to discrete pedagogical categories. In the second stage, a Large Language Model (LLM) equipped with Retrieval-Augmented Generation (RAG) converts each abstract plan step into a concrete, resource-grounded learning activity. An ensemble of five student behavior simulators supplies uncertainty-aware quality scores, and a Diffuser-based reinforcement learning module closes the optimization loop. Evaluated on two public educational datasets, DLPP yields 7–8% PKG improvement over the pure diffusion baseline on both EdNet and Junyi Academy. On the medium-scale EdNet (2941 training paths), it outperforms RL-Path and DiffPath baselines while remaining competitive with rule-based methods.
This paper introduces an LLM-Augmented Reinforcement Learning Agent that integrates LLM-driven planning with RL-based action optimization, and highlights a promising direction for building more capable autonomous systems.
Christophe D. Hounwanou, John Emeka Eze, Yaé Ulrich Gaba· 0 citations
Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is...
An execution-grounded dual-path consequence-aware agent for CLI-based SONiC operations, which generates multiple complete actions, predicts their execution consequences, and selects the final action through utility- and risk-aware reranking is proposed.
Yuxuan Chen, Rong-Peng Li, Zhi-Feng Zhao et al.· 0 citations
Deep research agents augment large language models with external tools to answer complex, long-horizon questions through multi-turn reasoning. Learning from prior experience is crucial for continual improvement, yet existing methods either retrieve verbose task-specific traces that burden decision-making, or distill pr...
Jie Ding, Rui Sun, Xin-Yi Zhang et al.· 1 citation
A coherent map of the rapidly expanding landscape of visual RL is provided to provide researchers and practitioners with a coherent map of the rapidly expanding landscape of visual RL and to highlight promising directions for future inquiry.
Value-Guided Flow Matching (VGFM), a scalable offline RL framework that enables dense value-guided shaping within a flow-based policy while avoiding BPTT and additional algorithmic overhead, is proposed.
Prajwal Koirala, Mark E. Campbell· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.