SPEAR: Phase-Aware Multi-Noise Removal for Supervised Fault-Event Detection and Localization in Simulated Factory Scenes
Abstract
Acoustic condition monitoring of factory machinery must operate under heavy multi-source interference: many machines run simultaneously, environmental noise is broadband, and workers speak while moving through the plant. A target machine fault is treated as a sustained acoustic event embedded in this multi-noise field, and the problem is cast as combined multi-noise removal, continuous fault-event detection, and source localization on a synchronized distributed microphone array. Unlike first-shot unsupervised anomaly detection, representative fault-event clips are assumed available at training time, so the task is treated as supervised target fault-event detection and localization. All experiments use simulated scenes; no real factory recordings are used. Taking the phase-aware lightweight framework DUAL-RES as a reference, a modified complex-mask denoiser with an integrated classifier is designed for the factory target-event setting, a multichannel spatial front-end that exploits inter-channel phase differences is introduced to provide spatial cues that help distinguish a directional fault from directional and diffuse interference, and the fault is localized with an event-mask-weighted, partially-whitened steered-response estimator, in which the learned event mask emphasizes the fault’s active frequency bands and partial whitening avoids the noise amplification of full phase transform (PHAT) weighting. On synthetic factory scenes rendered with room acoustics, DCASE machine sounds, DEMAND environmental noise, and moving or stationary speech, and over an operating range of −5 to + 10 dB event signal-to-noise ratio (event-SNR), the proposed framework, SPEAR (Spatial Phase-aware Event Acoustic Ranging), attains an operating-range event-detection area-under-curve (AUC) of 0.981 against 0.932 for an unsupervised Mahalanobis reference over the same range; the main advantage is at −5 dB, where the spatial front-end raises the AUC from 0.860 for a supervised single-channel model to 0.938 (and from 0.784 for the unsupervised reference). Including the below-range −10 dB condition, the full-range AUC is 0.945 against 0.859 for the unsupervised reference. For localization, the same steered-response estimator is weighted by the SPEAR event mask; the mask’s benefit is concentrated at low SNR, cutting the median horizontal error at −5 dB from 6.57 m for the unweighted estimator to 1.93 m (a $3.4\times $ improvement) and roughly halving the operating-range 90th-percentile error (14.68 to 8.45 m). Over 0 to + 10 dB both the unweighted partially-whitened estimator and SPEAR remain median sub-meter, with partial whitening rather than the mask the dominant factor; there the mask is slightly detrimental in median and accuracy within one meter, and a clear gap to the oracle-separation reference remains, especially at 0 dB. Below the operating range (−10 dB) the fault is globally about 10 dB below the aggregate background and provides too few event-dominant frequency bands for reliable mixture-based localization, a practical limit that is analyzed and left as future work. A hard-negative analysis further shows that the detector ranks speech and impulsive-knock transients below true faults but does not distinguish normal machine foregrounds, confirming that the task is supervised target-event detection rather than general anomaly discrimination. All learned components remain compact, with at most 1.04 M parameters, making SPEAR a candidate for future on-device deployment.