Interpretable Shared‐Backbone MobileViT Framework for sEMG‐Based Hand Motion Pattern Recognition and Finger Joint Angle Prediction
Abstract
Surface electromyography (sEMG) signals are vital bioelectrical indicators for decoding human motion intentions in rehabilitation robotics. However, achieving high‐precision synchronous modeling of hand motion pattern recognition and joint angle prediction remains challenging. This study proposes a shared‐backbone sEMG‐based modeling framework built upon the MobileViT backbone, which integrates the local feature extraction capability of convolutional structures with the global dependency modeling of Transformers. An improved energy kernel method for the two‐dimensional (2D) characterization of temporal data is developed to generate interpretable 2D characterization of muscle activity through matrix‐based quantification of amplitude–velocity distributions. In addition, a feature discriminant analysis approach is introduced to visualize and evaluate the separability and intra‐class compactness of deep features across model layers. Using the NinaPro DB7 dataset, two task‐specific models were constructed: MobileViT‐PR for motion pattern recognition and MobileViT‐MP for finger joint angle prediction. Experimental results demonstrate that the proposed framework achieved recognition accuracies of 99.8% and 99.7% for intact and amputee subjects, respectively. In regression tasks, the MobileViT‐MP model yielded an average coefficient of determination R‐Square of 0.925 and a root mean square error of 6.444. Furthermore, the effects of multiple factors—including window length, stride, exercise type, subject age, and condition—were systematically analyzed to provide comprehensive references for future sEMG modeling and rehabilitation studies. Overall, the proposed framework enables dual‐task modeling with high interpretability, discriminative capability, and computational efficiency, offering an effective solution for human–robot interaction and rehabilitation control.