Skip to content
Conference

WhisperEngineClassifier: A Transfer Learning Framework for Audio-Based Engine State Classification

Aug 2026 · 2026 6th International Conference on Emerging Smart Technologies and Applications (eSmarTA) · pp. 1-8 · 0 citations · 28 references

Abstract

The process of diagnosing and monitoring automotive systems through audio-based engine state classification has become more vital for intelligent vehicle systems as real-world applications face challenges from background noise, device differences, and class distribution problems which are issues addressed in this research. The VGG-Sound engine sound database was developed through ontology-based filtering, manual annotation, preprocessing, and five-state source-disjoint splitting. Eight CNN baselines were benchmarked under a unified training setup, revealing limitations in modeling long-range temporal dependencies. The development of WhisperEngineClassifier requires the adaptation of a pre-trained Whisper speech encoder through the removal of its decoder and the addition of lightweight pooling heads which will be fine-tuned in two different phases. The proposed model achieved 85.0% accuracy and 0.850 F1-score, outperforming the best CNN baseline, DenseNet169, by 16.2% in accuracy and showing clear gains on difficult classes. The system shows its ability to operate in real-world applications through a Flask-based deployment that uses FP16 quantization and GPU acceleration to support automotive telematics and predictive maintenance.

View source

Similar papers

Open access 2026

A System-Level Analysis of Cross-Domain Generalization Failure in Audio Deepfake Detection

Audio deepfake detectors often report high accuracy on individual benchmarks, yet their reliability under domain shift remains largely untested. This study presents a system-level analysis of cross-domain generalization failure, evaluating five hybrid architectures (CNN–LSTM, TCN, TCN–LSTM, Conformer, and TCM-Conformer...

Saadin Oyucu, Bilgehan Arslan, Şeref Sağıroğlu · 0 citations
Review Open access 2023

Deep Learning Approaches for Speech Recognition Systems

Automatic Speech Recognition (ASR) has evolved from rule-based and statistical methods to deep learning approaches, achieving near human-level performance under certain conditions. Traditional HMM-GMM models face limitations in handling long-term dependencies, speaker variability, and noise. Modern architectures such a...

Lydia Languish · 0 citations
Open access Sep 2026

AI4TEN: Fine-Tuning Pre-Trained Audio Transformers (BEATs, AST) for Cross-Dataset Acoustic Vehicle Classification with Domain Adaptation

Both pre-trained audio transformers substantially outperform the baseline for cross-dataset vehicle classification into five categories, with self-supervised pre-training producing slightly more transferable representations than supervised pre-training.

Jibran Khan · 0 citations
Open access Aug 2026

Generalization Analysis of YAMNet-DNN Architectures in Deepfake Audio Classification

The rapid advancement of speech synthesis and voice conversion technologies has increased the risk of audio deepfake attacks, necessitating robust and generalizable detection systems. This study proposes a deepfake audio detection framework that leverages pretrained YAMNet embeddings as a feature extractor, combined wi...

Hakam Dzakwan Diash, Dwi Arman Prasetya, Alfan Rizaldy Pratama et al. · 0 citations
Open access Sep 2026

Audio Deepfake Detection Using Dual-Branch CNN with Shared Weights

The dual-branch CNN with shared weights architecture presented here is augmented with self-attention modules to detect audio deepfakes with greater efficiency and improves feature extraction compared with the standard method.

Zainab A. Jawad, Ahmed J. Obaid · 0 citations
Review Open access Sep 2026

VoiceFusionNet: A Hybrid CNN--Transformer Framework for Speech-Based Parkinson’s Disease Screening Using Speech Signal Analysis

VoiceFusionNet, a hybrid CNN–Transformer framework evaluated on the Spanish PC-GITA Speech Corpus, applies amplitude normalization, short-time Fourier transforms, Mel-scale filtering, and MFCC extraction before learning local acoustic and global contextual representations through convolutional and Transformer modules.

V. V, A. Dumka · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.