UBG-Net features a Modality Uncertainty-aware Bayesian Fusion mechanism that injects signal-level aleatoric uncertainty into a Bayesian network to model epistemic uncertainty, thereby ensuring robust fusion of pre-trained backbone features.
Abstract
Audio-Visual speech recognition systems often degrade in real-world scenarios due to signal corruption and distribution shifts. To address this, we propose a unified uncertainty-modeling framework, namely the uncertainty-aware Bayesian gating network (UBG-Net). UBG-Net features a Modality Uncertainty-aware Bayesian Fusion (MUBF) mechanism that injects signal-level aleatoric uncertainty into a Bayesian network to model epistemic uncertainty, thereby ensuring robust fusion of pre-trained backbone features. For inference, we introduce Distribution Uncertainty-aware Hierarchical Voting (DUHV) to select transcripts from Monte Carlo samples, prioritizing frequency and using inference scores in case of a tie. Experiments on the AVCocktail and LRS2 datasets demonstrate the overall superiority of UBG-Net compared to SOTA baselines. Ablation studies confirm that MUBF and DUHV effectively filter noise, enhancing fusion and decoding robustness.
Detecting audio-visual DeepFake (AVDeepFake) is becoming increasingly important as synthetic media tools become widely accessible and spread across consumer devices. In this study, we present an Detecting audio-visual DeepFakes (AV-DeepFakes) has become increasingly critical with the rapid proliferation of accessible s...
Nasir Saleem, Ahmad Ali, Zhuo-Qi Zeng et al.· International Journal of Int...· 0 citations
Speaker verification back-ends commonly combine similarity scoring, score normalization, and calibration. However, speaker embeddings extracted from real-world utterances have trial-dependent reliability because of factors such as duration, noise, and channel variation. Existing uncertainty-aware methods primarily impr...
Speech enhancement models often struggle to adapt to heterogeneous and ambiguous noise conditions. This paper proposes a location-aware mixture-of-experts framework that jointly uses acoustic observations and categorical location context for adaptive speech enhancement. A gating network estimates continuous expert weig...
Large language model (LLM)-based audio-visual speech recognition (AVSR) systems are robust under noise. Contrastive decoding (CD), originally introduced to stabilize LLM generation by contrasting a weaker model against a stronger one at inference time, adjusts predictions without additional training. In this work, we a...
Youngchae Kim, Dali Yang, Joon-Hyuk Chang· 0 citations
With the development of wireless communications, the electromagnetic environment has become increasingly complex. Reliable signal recognition is fundamental for spectrum awareness and intelligent management. However, single-domain features are vulnerable to noise and fading, leading to limited performance. To address t...
Xuan-Di Qiao, Han Zhang, Yi Jiang et al.· 2026 IEEE/CIC International...· 0 citations
DSF-Net is a novel neural network that introduces a dual-strategy fusion approach for the AVSELD task, designed for robust and computationally efficient multi-modal comprehension.