Evaluating Self-Supervised Pretraining for Wrist-Worn Physiological Signals in Cardiac Estimation and Gesture Recognition
Abstract
Wearable systems that co-locate electrocardiography (ECG) and electromyography (EMG) on a single platform enable simultaneous cardiovascular monitoring and gesture recognition. Yet, developing robust models for such systems remains challenging due to limited labeled data, low channel counts, and motion-induced artifacts. We present an exploratory study comparing two representation paradigms for a custom dual-channel ECG-EMG wearable platform: (i) self-supervised learning (SSL) via masked autoencoders pretrained on large public biosignal datasets, and (ii) digital signal processing- (DSP-) based pipelines, combining with downstream classifiers for target applications. Using data collected from 20 subjects performing nine dynamic hand gestures, we evaluate both approaches on cardiac parameter estimation and gesture classification. Our findings indicate that the optimal representation paradigm is task-dependent. For cardiac parameter estimation, DSP-based detection on wrist ECG achieves higher heart rate estimation accuracy, while an SSL-based approach yields more motion-robust heart rate variability (HRV) estimation under motion. For gesture recognition, SSL embeddings substantially outperform DSP-based time frequency representations when paired with CNN classifiers, while remaining comparable under traditional machine learning classifiers, highlighting that SSL pretraining provides a more accurate and classifier-invariant representation for low-channel EMG.