Skip to content
Open access

Intelligent optimization of personalized learning path based on transformer and reinforcement learning.

Jul 2026 · Scientific Reports · 0 citations
Medicine

TL;DR

Findings suggest that closed-loop coordination among state representation, policy optimization, and educational constraints contributes to improvements in personalized learning path recommendation quality.

Abstract

Smart education platforms accumulate large volumes of learning resources and interaction logs. However, fixed recommendation sequences are often unable to adapt to differences in knowledge foundations, learning behaviors, and cognitive progress among learners. Existing learning path recommendation methods primarily focus on static matching and single-step ranking, while providing limited support for modeling learning state evolution and long-term path benefits. To improve learning resource recommendation accuracy, path continuity, and learning completion outcomes, a personalized learning path (PLP) optimization model integrating Transformer and Reinforcement Learning is developed. The model aims to generate learning resource sequences that align with individual cognitive processes through dynamic state awareness, sequential policy decision-making, and the incorporation of educational constraints. First, learning behavior sequences, knowledge point information, answer feedback, resource access records, and time intervals are encoded within a unified framework. A Transformer is employed to capture long-range dependencies and dynamic state features. Second, PLP generation is formulated as a continuous decision-making process, in which RL performs policy search within the candidate resource space. Finally, knowledge prerequisite relationships, difficulty progression rules, learning load boundaries, and a multi-objective reward function are incorporated into a unified optimization framework. Offline simulation experiments are conducted on three public datasets, namely EdNet, ASSISTments, and Junyi. Performance is compared with Self-Attentive Knowledge Tracing, Separable Self-Attentive Neural Knowledge Tracing Plus, Unified Knowledge Tracing (UniKT), and Time-Aware Reinforcement Learning. Experimental results indicate that Transformer and Reinforcement Learning for Learning Path (TFRL-Path) achieves a hit ratio of 0.846 on EdNet, exceeding UniKT by 0.025. On ASSISTments, the mean average precision reaches 0.737 and the completion rate reaches 0.801. On Junyi, precision reaches 0.736 and convergence epochs decrease to 38. The cross-dataset average composite score reaches 0.798 with a standard deviation of 0.028. The prerequisite satisfaction rate reaches 0.872, while the resource repetition rate decreases to 0.121. These findings suggest that closed-loop coordination among state representation, policy optimization, and educational constraints contributes to improvements in personalized learning path recommendation quality.

Read PDF

Similar papers

Open access Aug 2026

Dynamic optimization algorithm for personalized Japanese learning path driven by reinforcement learning

The framework successfully addresses the limitations of traditional rule-based and static systems by introducing a scalable, data-driven approach that adapts to individual learner needs, enhances engagement, supports adaptive personalized learning feedback, and continuously evolves to improve personalized learning outc...

Shu-Yun Bian · 0 citations
Conference Aug 2026

Recommendation of adaptive learning pathways for college Chinese based on reinforcement learning

This paper introduces a technically advanced recommendation framework for university-level Chinese language courses, which models student knowledge states and behavioral data as a Markov Decision Process and applies a deep Q-network to predict optimal content sequencing.

Liqun Fang · 0 citations
Conference Jul 2026

DLPP: a generative learning path planning framework based on LLM semantic guidance and closed-loop optimization

DLPP (Diffusion-based Learning Path Planning Framework), which recasts path generation as a two-stage process of instructional planning followed by semantic instantiation, outperforms RL-Path and DiffPath baselines while remaining competitive with rule-based methods.

Yuan Ren, Zhanfang Chen, Zeming Du et al. · 0 citations
Conference Aug 2026

Recommendation of adaptive learning paths for English MOOCs based on reinforcement learning

The findings demonstrate the effectiveness and scalability of combining DQN-based reinforcement learning with advanced student profiling in a MOOC environment and the model's resilience to noise and incomplete data.

Xin Zhang, Mei Li, Yujiao Han · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.