A transformer-based reinforcement learning framework for autonomous training of industrial robotic manipulators
Abstract
Autonomous robotic manipulation in modern manufacturing requires control policies that simultaneously achieve precision, adaptation, operational safety, and multi-task capability. This paper proposes a Transformer-Based Reinforcement Learning Network (TRL-Net) that combines temporal state encoding, task-conditioned representation, actor-critic reinforcement learning, domain-randomized simulation, and safety-constrained action projection. The evaluation is limited to ten simulated industrial manipulation tasks. The reported point estimates are a 98.6% task success rate, 0.81 mm average positioning error, 0.0% collision rate in the evaluated simulation episodes, 0.42 m/s³ trajectory jerk, and 4.2 ms policy inference time. These findings are consistent with improved temporal control and multi-task performance relative to the reported baselines. However, the archived experimental record does not contain real-robot trials, component-wise ablations, seed-level dispersion, complete hyperparameters, or hardware specifications. The conclusions are therefore restricted to the reported simulation setting, and the safety layer is interpreted as enforcing configured simulator constraints rather than guaranteeing safe operation on physical industrial hardware.