Aug 2026· 2026 6th International Conference on Emerging Smart Technologies and Applications (eSmarTA)· pp. 1-8· 0 citations· 28 references
Abstract
The process of diagnosing and monitoring automotive systems through audio-based engine state classification has become more vital for intelligent vehicle systems as real-world applications face challenges from background noise, device differences, and class distribution problems which are issues addressed in this research. The VGG-Sound engine sound database was developed through ontology-based filtering, manual annotation, preprocessing, and five-state source-disjoint splitting. Eight CNN baselines were benchmarked under a unified training setup, revealing limitations in modeling long-range temporal dependencies. The development of WhisperEngineClassifier requires the adaptation of a pre-trained Whisper speech encoder through the removal of its decoder and the addition of lightweight pooling heads which will be fine-tuned in two different phases. The proposed model achieved 85.0% accuracy and 0.850 F1-score, outperforming the best CNN baseline, DenseNet169, by 16.2% in accuracy and showing clear gains on difficult classes. The system shows its ability to operate in real-world applications through a Flask-based deployment that uses FP16 quantization and GPU acceleration to support automotive telematics and predictive maintenance.
Audio deepfake detectors often report high accuracy on individual benchmarks, yet their reliability under domain shift remains largely untested. This study presents a system-level analysis of cross-domain generalization failure, evaluating five hybrid architectures (CNN–LSTM, TCN, TCN–LSTM, Conformer, and TCM-Conformer...
Automatic Speech Recognition (ASR) has evolved from rule-based and statistical methods to deep learning approaches, achieving near human-level performance under certain conditions. Traditional HMM-GMM models face limitations in handling long-term dependencies, speaker variability, and noise. Modern architectures such a...
Lydia Languish· International Journal of App...· 0 citations
Both pre-trained audio transformers substantially outperform the baseline for cross-dataset vehicle classification into five categories, with self-supervised pre-training producing slightly more transferable representations than supervised pre-training.
The rapid advancement of speech synthesis and voice conversion technologies has increased the risk of audio deepfake attacks, necessitating robust and generalizable detection systems. This study proposes a deepfake audio detection framework that leverages pretrained YAMNet embeddings as a feature extractor, combined wi...
Hakam Dzakwan Diash, Dwi Arman Prasetya, Alfan Rizaldy Pratama et al.· International Journal of Adv...· 0 citations
The dual-branch CNN with shared weights architecture presented here is augmented with self-attention modules to detect audio deepfakes with greater efficiency and improves feature extraction compared with the standard method.
Zainab A. Jawad, Ahmed J. Obaid· Applied Informatics· 0 citations
VoiceFusionNet, a hybrid CNN–Transformer framework evaluated on the Spanish PC-GITA Speech Corpus, applies amplitude normalization, short-time Fourier transforms, Mel-scale filtering, and MFCC extraction before learning local acoustic and global contextual representations through convolutional and Transformer modules.
V. V, A. Dumka· VFAST Transactions on Softwa...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.