This work estimates time- and frequency-dependent power proportions by a causal low-complexity Deep Neural Network, from which multiple downstream acoustic scene analysis measures can in principle be derived by simple post-processing.
Abstract
Acoustic scene analysis is essential for adapting hearing-aid signal processing algorithms to the current listening environment. However, state-of-the-art (SOTA) systems typically rely on multiple independent estimators for tasks such as scene classification, Voice Activity Detection (VAD), or Signal-to-Noise Ratio estimation, which increases computational complexity and fails to exploit dependencies between related tasks. To address this problem, we propose a unified and interpretable acoustic scene representation by decomposing the observed mixture spectrum into speech, music, and noise power components. This is motivated by the typical listening targets of hearing-aid users. In particular, we estimate time- and frequency-dependent power proportions by a causal low-complexity Deep Neural Network, from which multiple downstream acoustic scene analysis measures can in principle be derived by simple post-processing. In this work, we validate the proposed representation using VAD as a representative downstream task and show performance comparable to a SOTA estimator while providing a substantially richer scene description.
In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual lea...
Debabrata Gogoi, Sushanta Kabir Dutta· Engineering Research Express· 0 citations
Voice Activity Detection (VAD) is a fundamental component that supports a wide range of audio/speech processing applications. Numerous studies have addressed VAD using a single microphone or a compact microphone array, yet their performance remains limited in distant scenarios. In this paper, we develop a novel data-dr...
De Hu, Shao-Jie Li, Qing-Ying Zhao et al.· IEEE Transactions on Audio,...· 0 citations
Hands-free communication devices and speakerphones are inherently affected by acoustic echo and background noise. To mitigate these impairments, end-to-end discriminatively trained neural networks have emerged as the best-performing approach in research and deployment. While recent advancements in generative methods ha...
Haljan Lugo, Ernst Seidel, Pejman Mowlaee et al.· 0 citations
Nowadays, microphone array (MA) technology has been installed into almost acoustic equipments, such as surveillance device, mobile phone, voice con trolled- device, teleconference system. Due to its high diversity, high directivity index towards the sound source, MA allows obtaining the original useful of desired talke...
Chau Thi Huyen Nguyen, Quan Trong The, Thang Minh Nguyen· Indonesian Journal of Electr...· 0 citations
Speech understanding in noise remains challenging for hearing-aid users, particularly in the presence of competing speakers. Conventional hearing aids typically perform speech enhancement (SE) and hearing-loss compensation in separate stages, which may cause enhancement errors and signal distortions to carry over to th...
Audio-visual speech enhancement (AVSE) aims at extracting target speech from multi-speaker mixtures by exploiting visual cues. Although recent studies have reported strong performance on simulated datasets, the performance, however, often drops dramatically when they are applied to real-world audio-visual recordings. T...
Tongtao Ling, Zhong-Qiu Wang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.