Skip to content
Open access

Multi-Resolution Acoustic Prototype Anchoring for Limited-Sample Industrial Anomalous Sound Detection

Jul 2026 · Energies · Vol 19, pp. 3530 · 0 citations · 17 references

TL;DR

Overall, TD-CPA provides a simple, normal-data-only, and interpretable solution for industrial anomalous sound detection under limited target-data conditions.

Abstract

Industrial anomalous sound detection is important for machine condition monitoring, especially when anomalous recordings are unavailable and only a small number of normal recordings can be collected from the target machine. This paper proposes target-domain cosine prototype anchoring (TD-CPA), a lightweight and interpretable framework for limited-sample industrial anomalous sound detection. The method combines multi-resolution Log-Mel representations with global, temporal-delta, and segment-wise statistical features to characterize both stable operating patterns and short-duration acoustic variations. PCA is then used to obtain compact embeddings, while ten target-normal recordings are employed to construct a target acoustic prototype. Anomaly scores are calculated based on the cosine deviation between each test embedding and this prototype. Experiments covering seven industrial machine categories show that TD-CPA achieves mean AUC, pAUC, and HScore values of 0.6067, 0.1389, and 0.2095, respectively, representing the highest numerical mean performance among the evaluated statistical, distance-based, compact autoencoder, and frozen pretrained-embedding baselines. Ablation experiments confirm the complementary contributions of the statistical feature components and multiple time-frequency resolutions. The target-sample sensitivity experiment further demonstrates that broader target-normal coverage generally improves detection performance. Overall, TD-CPA provides a simple, normal-data-only, and interpretable solution for industrial anomalous sound detection under limited target-data conditions.

Read PDF

Similar papers

Open access Aug 2026

Analysis of time–frequency signal representations for anomaly detection in industrial environments

In the context of Industry 4.0, sound-based anomaly detection has emerged as a relevant strategy to support predictive maintenance and reduce unplanned downtime in industrial environments. This study investigates the impact of different time–frequency representations on acoustic anomaly detection in industrial machines...

P. D. de Araújo, Rodrigo de Paula Monteiro, C. Bastos · 0 citations
Preprint Aug 2026

Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning

In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farther away to capture noise. For this task,...

Takuya Fujimura, G. Wichern, Yoshiki Masuyama et al. · 0 citations
Open access Oct 2026

Fusion-Inception-Aux: An Auxiliary-Supervised Multi-Scale Fusion Network for Distributed Acoustic Sensing Event Recognition

Distributed acoustic sensing (DAS) event recognition is challenging because disturbance events may have similar temporal waveforms, heterogeneous channel responses, and strict false-alarm requirements in practical monitoring systems. This paper proposes <bold>Fusion-Inception-Aux</bold>, an auxiliary-supervised dual-br...

Tian-Chang Xie, Hai-Ling Wang, Wei-Guang Wang et al. · 0 citations
Preprint Sep 2026

Domain-Incremental Learning for Multi-Channel Replay Speech Detection

Replay attacks are the most accessible threat to voice-controlled systems, and the acoustic cues that expose them are strongly modulated by the environment in which the attack is mounted. A detector deployed in the field therefore has to absorb new acoustic conditions over time, ideally without revisiting past recordin...

Michael Neri · 0 citations
Open access Aug 2026

A Hybrid MFCC–WavLM Feature Fusion for Audio Deepfake Detection

Recent advances in generative artificial intelligence have enabled highly realistic speech synthesis using text-to-speech (TTS), voice conversion (VC), and neural voice cloning techniques, posing significant security threats to Automatic Speaker Verification (ASV) systems. Conventional handcrafted features such as Mel-...

K. S. Kumar, Madduluri Suneetha, K. R. Anudeep Laxmi Kanth et al. · 0 citations
Open access Aug 2026

Dual-Channel Acoustic Temporal–Spectral Representation and Noise-Perturbation Learning for Mining Conveyor Fault Diagnosis

This study develops a robust acoustic fault diagnosis framework for mining conveyor idlers, addressing the challenge of detecting early-stage mechanical degradation in noisy and imbalanced industrial environments. A dual-channel temporal--spectral representation is constructed by combining multi-scale log-Mel spectrogr...

Zhenyu Wang, Tai-Li Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.