Predefined-Time Nonsingular Sliding Mode Control for Markov Jump Manipulator Systems via Reinforcement Learning
Abstract
This paper investigates the trajectory-tracking control problem for manipulators subject to stochastic payload variations and composite disturbances. The manipulator dynamics are modeled as a Markov jump system to capture random mode switching. To ensure high-precision tracking, a novel reinforcement learning (RL)-based predefined-time nonsingular sliding mode control (PNTSMC) strategy is proposed. Specifically, a new sliding mode surface is designed to guarantee nonsingular convergence within a predefined time independent of the initial conditions. Furthermore, a radial basis function (RBF) neural-network Actor-Critic RL framework driven by instantaneous performance is developed to compensate for composite disturbances online. The practical predefined-time stability of the closed-loop system is rigorously established via Lyapunov analysis. Simulation results confirm that the tracking error converges to a small residual neighborhood within a predefined time despite Markovian mode switching. These results verify that the PNTSMC with RL-based disturbance compensation scheme provides a robust and high-precision solution for manipulators in complex dynamic environments. Note to Practitioners—This paper addresses high-precision trajectory tracking for manipulators subject to stochastic payload variations and composite disturbances. Traditional sliding mode control (SMC) methods often struggle with control signal singularities and convergence times that depend on initial system states, making it difficult to guarantee precise execution times. To overcome these practical bottlenecks, this article proposes a novel framework integrating PNTSMC with an Actor-Critic reinforcement learning-based disturbance compensation mechanism. By modeling the dynamics as a Markov jump system, the scheme ensures nonsingular convergence within a predefined time independent of the initial conditions, providing a robust solution for operations in complex dynamic environments.