Benchmarking Multi-Agent Reinforcement Learning for Stochastic OSAT Scheduling: A Reproducible Computational Study for Semiconductor Operations Management
TL;DR
The results therefore support selective, KPI-specific learned-policy benefits rather than universal MARL superiority and provides a basis for longer-horizon, multi-seed, factory-calibrated, and hybrid RL-heuristic validation in semiconductor operations management.
Abstract
Outsourced Semiconductor Assembly and Test (OSAT) operations require repeated scheduling decisions across interdependent stages under processing-time variability, equipment failures, product-mix changes, and demand uncertainty. This study evaluates whether alternative multi-agent reinforcement-learning (MARL) coordination structures provide measurable short-horizon scheduling benefits when compared in a common stochastic OSAT environment. Four learned policy structures—H-MARL-CTDE, H-MARL-Independent, H-MARL-Value, and Flat-MARL—were evaluated against FIFO, SPT, and EDD. The executable simulator represents Die Attach, Wire Bond, Encapsulation, Final Test, and Quality Control, advances in five-minute steps over six simulated hours, and uses active stochastic arrivals, product sampling, lognormal processing-time variation, and MTBF/MTTR-based failure and repair dynamics. The final matched design comprises 27 variability-demand-failure scenarios, ten replications, seven policies, and 1,890 run-level observations. Policy differences were assessed using matched-block Friedman tests and paired Wilcoxon signed-rank tests with Holm correction. H-MARL-Independent achieved the highest mean throughput (16.4932 units/hour) and the lowest energy per completed unit (3.8375), while H-MARL-CTDE achieved the highest OEE proxy (53.690%) and a mean flow time of 0.5848 h. SPT achieved the lowest flow time (0.5838 h). Global policy differences were significant for throughput, flow time, energy efficiency, and OEE (all p<.001), but not makespan (p=.987). CTDE and Independent significantly outperformed FIFO on the four discriminating KPIs after Holm correction, but neither consistently outperformed SPT. The results therefore support selective, KPI-specific learned-policy benefits rather than universal MARL superiority. The primary contribution is an auditable common benchmark that separates computational evidence from production claims and provides a basis for longer-horizon, multi-seed, factory-calibrated, and hybrid RL-heuristic validation in semiconductor operations management.