Skip to content
Review Open access

An Evidence-Based Systematic Literature Review of Deep Reinforcement Learning for Manufacturing Scheduling

Sep 2026 · Mathematics · Vol 14, pp. 3280 · 0 citations · 84 references

TL;DR

An evidence-based systematic literature review of DRL for manufacturing scheduling using a structured methodology for corpus construction, configuration-level coding, evidence traceability, study-quality assessment, and cross-study synthesis is presented.

Abstract

Deep reinforcement learning (DRL) has become an important approach for manufacturing scheduling because it supports sequential decision-making under complex and changing production conditions. However, existing reviews primarily organize the literature by scheduling problem or learning method, providing less explicit support for tracing how manufacturing context, Markov Decision Process (MDP) formulation, scheduler architecture, and evaluation choices interact across heterogeneous studies. This study presents an evidence-based systematic literature review of DRL for manufacturing scheduling using a structured methodology for corpus construction, configuration-level coding, evidence traceability, study-quality assessment, and cross-study synthesis. The validated corpus comprises 52 primary studies and 54 independently coded DRL configurations. The evidence is synthesized across manufacturing scheduling characteristics, MDP design, DRL scheduler design, hybrid optimization, and empirical evaluation. The results show that scheduler design is context-dependent and architecturally diverse: manufacturing requirements are associated with differences in state, action, and reward formulation, while DRL schedulers combine different learning algorithms, representation architectures, control structures, and complementary optimization mechanisms. The evidence does not establish universal superiority for individual representations, algorithms, or hybrid architectures because reported outcomes remain strongly conditioned by problem formulation and experimental design. Evaluation evidence further highlights limited generalization, uneven statistical and component-level validation, and a continuing gap between benchmark or simulation studies and live industrial deployment. By linking study-, configuration-, and evidence-level information, this review provides a traceable basis for interpreting methodological relationships, identifying research gaps, and guiding the development and evaluation of DRL-based manufacturing scheduling systems.

Read PDF

Similar papers

Conference Open access Sep 2026

Preference-Guided Multi-Policy Optimization for Flexible Job Shop Scheduling

PGMPO is proposed, a novel learning framework consisting of a simple but effective multi-policy modeling approach that allows a single network to represent multiple decision-makers, and a preference-driven model optimization method that effectively guides policies to learn diverse and specialized problem-solving strate...

Inguk Choi, Woo-Jin Shin, Sang-Hyun Cho et al. · 0 citations
#machine learning Review Sep 2026

Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap

The growing demand for real-time, data-driven decision-making in complex and dynamic systems is placing increasing pressure on traditional Operational Research (OR) methodologies. Reinforcement learning (RL) has emerged as a complementary approach, offering strong learning and computational capabilities for sequential...

Ya-Han Lu, Dong-Yang Xia, Nurşen Aydın et al. · 0 citations
Open access Sep 2026

A hierarchical reinforcement learning based subtask-coordinated scheduling method for constrained multi-objective evolutionary algorithm

In recent years, constrained multi-objective optimization problems(CMOPs) remain challenging due to the complex structure of feasible regions, the difficulty of balancing convergence and diversity, and the lack of adaptive operator scheduling mechanisms. To address these issues, this paper proposes a hierarchical reinf...

Lu-Peng Hao, Yahui Shan, Guangyin Jin et al. · 0 citations
Review Open access Aug 2026

A survey on LLM-enhanced reinforcement learning in financial markets

A three-paradigm taxonomy (feature-based, auxiliary-based, and policy-based) based on the functional role of LLMs within the RL pipeline is proposed, which provides superior scalability and stability, though often at the expense of representational depth.

Ghusoon Hadi al-Aldaffaie, Alireza Taheri, Amirfarhad Farhadi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.