Open access
2026
Benchmarking Multi-Agent Reinforcement Learning for Stochastic OSAT Scheduling: A Reproducible Computational Study for Semiconductor Operations Management
The results therefore support selective, KPI-specific learned-policy benefits rather than universal MARL superiority and provides a basis for longer-horizon, multi-seed, factory-calibrated, and hybrid RL-heuristic validation in semiconductor operations management.
Mai Ngoc Huy
· International journal of res... · 0 citations