Skip to content
Open access

Benchmarking Multi-Agent Reinforcement Learning for Stochastic OSAT Scheduling: A Reproducible Computational Study for Semiconductor Operations Management

2026 · International journal of research and innovation in social science · Vol 10, pp. 3058-3066 · 0 citations

TL;DR

The results therefore support selective, KPI-specific learned-policy benefits rather than universal MARL superiority and provides a basis for longer-horizon, multi-seed, factory-calibrated, and hybrid RL-heuristic validation in semiconductor operations management.

Abstract

Outsourced Semiconductor Assembly and Test (OSAT) operations require repeated scheduling decisions across interdependent stages under processing-time variability, equipment failures, product-mix changes, and demand uncertainty. This study evaluates whether alternative multi-agent reinforcement-learning (MARL) coordination structures provide measurable short-horizon scheduling benefits when compared in a common stochastic OSAT environment. Four learned policy structures—H-MARL-CTDE, H-MARL-Independent, H-MARL-Value, and Flat-MARL—were evaluated against FIFO, SPT, and EDD. The executable simulator represents Die Attach, Wire Bond, Encapsulation, Final Test, and Quality Control, advances in five-minute steps over six simulated hours, and uses active stochastic arrivals, product sampling, lognormal processing-time variation, and MTBF/MTTR-based failure and repair dynamics. The final matched design comprises 27 variability-demand-failure scenarios, ten replications, seven policies, and 1,890 run-level observations. Policy differences were assessed using matched-block Friedman tests and paired Wilcoxon signed-rank tests with Holm correction. H-MARL-Independent achieved the highest mean throughput (16.4932 units/hour) and the lowest energy per completed unit (3.8375), while H-MARL-CTDE achieved the highest OEE proxy (53.690%) and a mean flow time of 0.5848 h. SPT achieved the lowest flow time (0.5838 h). Global policy differences were significant for throughput, flow time, energy efficiency, and OEE (all p<.001), but not makespan (p=.987). CTDE and Independent significantly outperformed FIFO on the four discriminating KPIs after Holm correction, but neither consistently outperformed SPT. The results therefore support selective, KPI-specific learned-policy benefits rather than universal MARL superiority. The primary contribution is an auditable common benchmark that separates computational evidence from production claims and provides a basis for longer-horizon, multi-seed, factory-calibrated, and hybrid RL-heuristic validation in semiconductor operations management.

Read PDF

Similar papers

#reinforcement learning Open access Sep 2026

Research on Multi-Objective Task Scheduling Optimization and Reinforcement Learning Decision-Making Methods for Edge-Cloud Collaborative Scalable Information Systems

INTRODUCTION: Edge-cloud schedulers must coordinate latency, energy, load balance, and deadline compliance under changing demand while keeping task-arrival and throughput units physically consistent.OBJECTIVE: This study evaluates MORL-ECSO under an auditable, paired-seed simulation protocol and compares it with tuned...

Li-Na Guo, Cheng-Yu Sun · 0 citations
#graph neural networks Open access Aug 2026

Multi-Agent Reinforcement Learning for Dynamic Inventory Rebalancing and Last-Mile Fulfillment Under Supply Chain Disruptions

Experiments show that prediction must be coupled with autonomous optimization to deliver practical resilience, and show that the learned policy sustains service levels, shortens recovery time, and reduces total disruption cost relative to base-stock and single-agent baselines.

Sohail Sayed, Nauman Sayed · 0 citations
Open access Sep 2026

Priority-Guided Action-Masked Proximal Policy Optimization for Dynamic Scheduling of Electric Power Material Verification Tasks Under Multi-Source Disturbances

The metering verification center of Yunnan Power Grid processes dynamically arriving batches of single-phase and three-phase smart meters, low-voltage current transformers, and collection terminals through multi-operation test chains on six automated verification lines under batch arrivals, urgent re-verification order...

Zhao-Lei He, Ao He, Cong Lin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Planning or Learning: Reliability and Cost in Multi-Asset Maintenance

Industrial maintenance systems involve multiple interacting assets and shared resources, making it challenging to balance reliability and operational cost using a single decision framework. While recent work has focused on reinforcement learning (RL) for maintenance scheduling, direct comparisons with planning approach...

X. Lee, Chandrasekar Venkatraman, Ahmed K. Farahat · 0 citations
Open access 2026

Entropy-Regulated Job-Shop Scheduling: A Bottom–Up Artificial Bee Colony Algorithm for Semiconductor Manufacturing

Analysis of queue-level dynamics reveals more regular behavior in the evaluated scenarios, with reduced fluctuations in queue lengths, batch waiting, and minimum queue entropy over time, indicating that the proposed ABC-based approach can improve observed predictability at the queue level.

Elnaz Khatmi, Khalil Al-Rahman Youssefi, Wilfried Elmenreich · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.