Skip to content
Preprint

A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

A hybrid CNN-SSM-Attention backbone for 12-lead ECG classification is introduced, which provides strong supervised baselines under a compact parameter budget and improves transfer, particularly in reduced-label settings and under both full fine-tuning and LoRA-based adaptation.

Abstract

Automatic 12-lead electrocardiogram (ECG) classification requires representations that jointly capture local waveform morphology, long-range temporal dynamics, and cross-lead dependencies, yet integrating these properties within a single efficient architecture remains challenging. This paper introduces a hybrid CNN-SSM-Attention backbone for 12-lead ECG classification. A convolutional stem performs early waveform tokenization and temporal reduction, mixed state-space and depthwise-convolutional blocks model temporal dynamics and local morphology, and a late self-attention stage enables global token interaction at reduced resolution. To improve transfer from unlabeled data, we further develop an ECG-oriented Joint-Embedding Predictive Pretraining (JEPA) framework. Unlike ViT-based JEPA methods that mask patch tokens before the encoder, the proposed method samples span masks at the latent temporal resolution and projects them back to the waveform domain, then predicts clean latent targets from a momentum encoder without waveform reconstruction. Experiments on CPSC2018, Chapman-Shaoxing, and PTB-XL, with pretraining on approximately 350K unlabeled CODE-15 recordings, show that the proposed backbone provides strong supervised baselines under a compact parameter budget. JEPA pretraining further improves transfer, particularly in reduced-label settings and under both full fine-tuning and LoRA-based adaptation. Code: https://github.com/yakoubbazi/Hybrid_ECG_Jepa

View source

Similar papers

Open access Sep 2026

A convolutional attention transformer network for ECG beat classification

This work aims to improve automated ECG beat classification in accordance with the Association for the Advancement of Medical Instrumentation (AAMI)-recommended classes by learning both local morphological patterns and longer-range temporal dependencies from ECG segments. We develop a complete analysis pipeline using t...

Shanmukha Rao Narsupalli, Rajesh Kumar Pullagura, Rajeswara Rao Gangula · 0 citations
#artificial intelligence Preprint Sep 2026

DR-net-Mamba: Selective State-Space Modeling for Long-Range ECG Time-Series Denoising

Electrocardiogram (ECG) recordings are corrupted by non-stationary noise sources that degrade diagnostic reliability, particularly in ambulatory and long-duration recordings. Deep learning denoisers exist, but convolutional architectures are limited by their receptive field, transformer-based models scale quadratically...

Basile Morel, Samuel Ruipérez-Campillo, Andreas P. Streich et al. · 0 citations
Aug 2026

HCFT: A Hierarchical Convolutional Fusion Transformer for Cross-Task EEG Decoding.

Electroencephalography (EEG) decoding remains challenging due to the non-stationary nature of neural signals and the limited generalization of existing models across tasks and subjects. To address this challenge, we propose a lightweight and generalizable decoding framework named Hierarchical Convolutional Fusion Trans...

Haodong Zhang, Jiapeng Zhu, Yitong Chen et al. · 0 citations
Open access Sep 2026

Higher-dimensional embedding of time-series data for machine learning

Deep learning has revolutionized image analysis, yet most clinical biosignals, especially multi-lead electrocardiograms (ECGs), remain one-dimensional and awkward for modern vision models. We introduce an orthogonal-polynomial imaging (OPI) framework that encodes 12-lead ECGs into a single two-dimensional im...

Karan Singh, Pingal Pratyush Nath, U. Sinha et al. · 0 citations
Open access Aug 2026

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al. · 0 citations
Book Open access Aug 2026

Bridging ECG and PPG: Latent-Space Prediction for Robust Physiological Analysis

Foundation models for physiological signals have shown promise for health monitoring, with recent work exploring multi-modal approaches that leverage complementary information across modalities such as PPG and ECG. However, existing methods typically rely on contrastive objectives or reconstruction losses that operate...

Zhaoliang Chen, Saurabh Kataria, Minxiao Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.