MTENet: A Multi-Representation Time-Series Evidential Network for Automated Heart Murmur Detection from Phonocardiogram Signals
Abstract
Early diagnosis is essential for the effective management of cardiovascular diseases (CVDs). Although conventional auscultation is the primary screening method, its reliance on subjective interpretation and susceptibility to clinical background noise have positioned phonocardiogram (PCG) analysis as a key diagnostic tool, making reliable automated interpretation a pressing necessity. Automated pipelines have evolved from handcrafted-feature machine learning to deep learning and transformer-based architectures, but the latter often depend on heavy time-frequency preprocessing and large parameter counts, inflating computational cost and increasing the risk of overfitting on limited or noisy clinical data. We propose MTENet, a Multi-representation Time-series Evidential Network that models a phase-enhanced one-dimensional PCG waveform through a bidirectional Mamba state-space encoder, capturing long-range temporal dependencies with linear-time complexity. Recordings are prepared by a label-independent, record-internal stage that combines an adaptive FFT filter bank with an automatically estimated cardiac-phase gain and returns a waveform of unchanged length and sampling rate, so that the model input remains a time-series rather than a fixed feature representation. The Mamba encoder forms one of three parallel branches, alongside an implicit neural representation (INR) branch for continuous signal modelling and a Mel-spectrogram branch computed on-the-fly within the network for spectral structure. The three streams are merged by a softmax-gated fusion. A multi-scale convolutional stem captures local transient structure, while the bidirectional Mamba encoder models longer-range cardiac rhythm. With approximately 2.58 million parameters, the model captures intricate temporal patterns while distributing representational responsibility across complementary streams. Under record-grouped four-fold validation on the primary HLS-CMDS corpus, MTENet attained an accuracy of 0.9816, a balanced accuracy of 0.9812, and an AUROC of 0.9938 under a leakage-free, record-level nested protocol in which the training epoch is selected on an inner-validation split drawn only from the training partition, so that the outer evaluation fold never informs model selection. Within-dataset evaluation on CirCor DigiScope 2022 yielded an AUROC of 0.9698.