Skip to content
Preprint

BiTSE: Binaural Target Speaker Extraction in Noisy Multi-Talker Environments for AR Glass Arrays

Aug 2026 · 0 citations · 28 references
Engineering

Abstract

Isolating a desired speech signal in noisy multi-talker conversational scenarios is a key requirement for augmented reality (AR) wearable microphone array systems. In this work, a binaural target speaker extraction (TSE) framework, termed BiTSE, is proposed. It leverages both spatial and temporal cues, specifically the direction-of-arrival (DoA) of the target speaker and corresponding voice activity information, to guide the extraction process. Built upon a binaural signal denoising architecture, our model integrates three key enhancements: (i) a DoA-aware attention mechanism using cyclic positional embeddings, (ii) a timestamp-based masking strategy that utilizes speaker activity to suppress non-target segments, and (iii) a novel two-stage loss optimization strategy that first trains the model for robust denoising and then fine-tunes it to improve perceptual quality. Evaluations on the SPeech Enhancement for Augmented Reality (SPEAR) challenge dataset demonstrate that the proposed BiTSE consistently improves upon conventional approaches, leading to enhanced signal fidelity and perceptual quality.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.