Skip to content
Open access

Preference-conditioned deep reinforcement learning for dynamic scheduling in sustainable and robust manufacturing

Aug 2026 · Production Engineering · Vol 20 · 0 citations · 39 references

TL;DR

The study demonstrates the potential of CPPS-oriented and preference-conditioned DRL for adaptive, energy-aware, and robust scheduling in smart manufacturing systems.

Abstract

Modern manufacturing requires scheduling methods that adapt to changing order arrivals, machine disruptions, customer priorities, stakeholder preferences, and time-varying energy conditions. This paper proposes a preference-conditioned deep reinforcement learning (DRL) approach for dynamic scheduling in sustainable and robust manufacturing. The approach is embedded in a cyber-physical production system (CPPS)-oriented framework that links production states, machine availability, energy-related background data, simulation-based learning, performance monitoring, and decision support. Within this framework, a Double Deep Q-Network (DDQN) scheduler is developed for joint job sequencing, machine assignment, and start-time adjustment. The scheduler uses a candidate-based state representation for dynamic order arrivals, vector-valued Q-output for objective-specific value estimation, and a priority- and preference-aware reward design. Customer priorities are treated as order-level attributes, while stakeholder preferences are encoded as system-level objective weightings. This enables one policy to consider energy-related cost, carbon emissions, energy demand, and tardiness while adapting to different preference profiles. The concept is demonstrated in an on-demand manufacturing (ODM)-oriented parallel CNC machining case with heterogeneous orders, product-specific setup and processing requirements, hourly electricity prices, carbon-intensity signals, and curriculum-adaptive machine breakdowns. DDQN is compared with three dispatching rules and two DRL baselines under shared training and testing scenarios. The results show that DDQN achieves the lowest energy-related cost and carbon emissions in training and unseen testing while maintaining acceptable delivery performance. Overall, the study demonstrates the potential of CPPS-oriented and preference-conditioned DRL for adaptive, energy-aware, and robust scheduling in smart manufacturing systems.

Read PDF

Similar papers

Review Open access Sep 2026

An Evidence-Based Systematic Literature Review of Deep Reinforcement Learning for Manufacturing Scheduling

An evidence-based systematic literature review of DRL for manufacturing scheduling using a structured methodology for corpus construction, configuration-level coding, evidence traceability, study-quality assessment, and cross-study synthesis is presented.

Yi-Kai Su, Chun-Jan Tseng · 0 citations
Open access 2026

Benchmarking Multi-Agent Reinforcement Learning for Stochastic OSAT Scheduling: A Reproducible Computational Study for Semiconductor Operations Management

The results therefore support selective, KPI-specific learned-policy benefits rather than universal MARL superiority and provides a basis for longer-horizon, multi-seed, factory-calibrated, and hybrid RL-heuristic validation in semiconductor operations management.

Mai Ngoc Huy · 0 citations
Open access 2026

Learning to Schedule Machines and Operators: A Human-Aware Deep Q-Learning Framework for Dynamic Flexible Job Shops

Most learning-based schedulers for job shop problems assume static workforce performance. However, dynamic dual-resource job shops must navigate stochastic disturbances and time-varying task durations due to human factors. This study proposes a human-centered Deep Reinforcement Learning framework, DD4LQN, for dynamic f...

Taji Hajar, Ayad Ghassane, Z. Abd-El-Hamid et al. · 0 citations
Conference Aug 2026

Data-Driven Real-Time Scheduling Method Based on Deep Reinforcement Learning for IIoT-Enabled Discrete Manufacturing Workshop

In a discrete manufacturing workshop with individualized production tasks and multi-disturbance production processes, real-time scheduling is urgently needed to shorten order completion time. Therefore, a real-time scheduling framework based on the industrial internet of things and deep reinforcement learning is propos...

Dao-Yuan Liu, Kai-Tuo Wang · 0 citations
Conference Open access Sep 2026

Preference-Guided Multi-Policy Optimization for Flexible Job Shop Scheduling

PGMPO is proposed, a novel learning framework consisting of a simple but effective multi-policy modeling approach that allows a single network to represent multiple decision-makers, and a preference-driven model optimization method that effectively guides policies to learn diverse and specialized problem-solving strate...

Inguk Choi, Woo-Jin Shin, Sang-Hyun Cho et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.