Skip to content
Preprint

DNN-Based Frequency-Dependent Estimation of Speech, Music, and Noise Power in Acoustic Mixtures for Hearing-Aid Scene Analysis

Aug 2026 · 0 citations · 25 references
Engineering

TL;DR

This work estimates time- and frequency-dependent power proportions by a causal low-complexity Deep Neural Network, from which multiple downstream acoustic scene analysis measures can in principle be derived by simple post-processing.

Abstract

Acoustic scene analysis is essential for adapting hearing-aid signal processing algorithms to the current listening environment. However, state-of-the-art (SOTA) systems typically rely on multiple independent estimators for tasks such as scene classification, Voice Activity Detection (VAD), or Signal-to-Noise Ratio estimation, which increases computational complexity and fails to exploit dependencies between related tasks. To address this problem, we propose a unified and interpretable acoustic scene representation by decomposing the observed mixture spectrum into speech, music, and noise power components. This is motivated by the typical listening targets of hearing-aid users. In particular, we estimate time- and frequency-dependent power proportions by a causal low-complexity Deep Neural Network, from which multiple downstream acoustic scene analysis measures can in principle be derived by simple post-processing. In this work, we validate the proposed representation using VAD as a representative downstream task and show performance comparable to a SOTA estimator while providing a substantially richer scene description.

View source

Similar papers

Open access Aug 2026

A deep residual complex learning framework with long-range temporal context for phase-aware speech enhancement

In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual lea...

Debabrata Gogoi, Sushanta Kabir Dutta · 0 citations
2026

Low-Rate Voice Activity Detector Over Wireless Acoustic Sensor Networks

Voice Activity Detection (VAD) is a fundamental component that supports a wide range of audio/speech processing applications. Numerous studies have addressed VAD using a single microphone or a compact microphone array, yet their performance remains limited in distant scenarios. In this paper, we develop a novel data-dr...

De Hu, Shao-Jie Li, Qing-Ying Zhao et al. · 0 citations
Preprint Sep 2026

DiffVQE2: An Efficient Low-delay Diffusion Model for Acoustic Echo and Noise Control

Hands-free communication devices and speakerphones are inherently affected by acoustic echo and background noise. To mitigate these impairments, end-to-end discriminatively trained neural networks have emerged as the best-performing approach in research and deployment. While recent advancements in generative methods ha...

Haljan Lugo, Ernst Seidel, Pejman Mowlaee et al. · 0 citations
Open access Sep 2026

An improved MVDR beamformer’s noise reduction based on post- filtering

Nowadays, microphone array (MA) technology has been installed into almost acoustic equipments, such as surveillance device, mobile phone, voice con trolled- device, teleconference system. Due to its high diversity, high directivity index towards the sound source, MA allows obtaining the original useful of desired talke...

Chau Thi Huyen Nguyen, Quan Trong The, Thang Minh Nguyen · 0 citations
Preprint Oct 2026

Audiovisual joint learning for end-to-end hearing aids

Speech understanding in noise remains challenging for hearing-aid users, particularly in the presence of competing speakers. Conventional hearing aids typically perform speech enhancement (SE) and hearing-loss compensation in separate stages, which may cause enhancement errors and signal distortions to carry over to th...

You-Jin Li, Yu Tsao, B. Su et al. · 0 citations
Preprint Aug 2026

Separate First, Then Associate: A Two-Stage Approach for Real-World Audio-Visual Speech Enhancement

Audio-visual speech enhancement (AVSE) aims at extracting target speech from multi-speaker mixtures by exploiting visual cues. Although recent studies have reported strong performance on simulated datasets, the performance, however, often drops dramatically when they are applied to real-world audio-visual recordings. T...

Tongtao Ling, Zhong-Qiu Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.