Sep 2026· Artificial Life and Robotics· 0 citations· 9 references
TL;DR
This framework provides a potential scalable pathway for domains with critically limited annotated data using massive synthetic data without synthetic labels to train a lightweight classifier on limited clinical data.
Abstract
Automated auscultation using wearable devices is essential for remote respiratory monitoring, but deep learning models often struggle to generalize due to the severe scarcity of annotated abnormal respiratory sounds. Since directly using synthetic data for supervised training risks learning artifacts instead of true pathological features, we propose a robust three-phase pipeline leveraging massive synthetic data without synthetic labels. First, a modified StyleGAN2 natively synthesizes rectangular Mel-spectrograms to preserve high-temporal-resolution acoustic characteristics, validated by kernel audio distance. Second, 100,000 synthetic spectrograms are used for unsupervised variational autoencoder pre-training, introducing a parallel asymmetric convolutional block to independently capture distinct time and frequency semantics. Finally, the encoder is repurposed as a feature extractor to train a lightweight classifier on limited clinical data. Empirical evaluations demonstrate that while a serial asymmetric kernel geometry yields the highest classification accuracy under frequency-dominant pathologies, our parallel architecture achieves highly competitive performance using only 43% of conventional square baseline parameters and rivals the massive pretrained audio neural networks CNN14 model using merely 0.5% of its convolutional footprint. This framework provides a potential scalable pathway for domains with critically limited annotated data.
Respiratory and cardiovascular diseases represent significant global health burdens, particularly in resource-constrained settings where access to specialized medical expertise is limited. This paper presents a comprehensive machine learning pipeline for auscultation-based diagnosis that converts short electronic steth...
Mohammed Mafaz Nadherssa· International Journal of Int...· 0 citations
Interstitial lung disease (ILD) screening from respiratory sounds (RSs) remains challenging due to the subtle acoustic differences between pathological and healthy patterns, compounded by limited labeled medical audio data. This paper proposes a curvature-controlled (CC) contextual–perceptual feature fusion framework f...
Ayushi Pal, Udit Satija, Jimson Mathew et al.· IEEE Signal Processing Lette...· 0 citations
Random or central placement of respiratory sounds within padded windows can cause length-based masks to exclude recorded audio. We developed an exact support-propagation interface for the pretrained bidirectional encoder representation from audio transformers (BEATs), mapping sample support through filterbank frames an...
Jie Niu, Peng Li· Medical Engineering and Phys...· 0 citations
This work proposes a framework that aligns self-supervised respiratory encoders with medical terminology in a shared latent space turning them into a zero-shot-capable foundation model, and uses a medical LLM to synthesize structured reports from metadata, creating dense semantic anchors for contrastive learning.
Mustafa Talha İlerisoy, Hung Manh Pham, Mathias Funk et al.· 0 citations
The rapid global spread of COVID-19 has highlighted the need for efficient and scalable respiratory disease screening methods. Compared with conventional diagnostic approaches such as RT-PCR, cough sound–based analysis provides a non-invasive and low-cost alternative. This paper proposes a lightweight deep learning fra...
Lixinyu Lu, Qifeng Han, Helu Hu et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.