Skip to content
Preprint

UBG-Net: An Uncertainty-aware Bayesian Gating Network for Robust Audio-Visual Speech Recognition

Jul 2026 · 0 citations · 30 references
Engineering

TL;DR

UBG-Net features a Modality Uncertainty-aware Bayesian Fusion mechanism that injects signal-level aleatoric uncertainty into a Bayesian network to model epistemic uncertainty, thereby ensuring robust fusion of pre-trained backbone features.

Abstract

Audio-Visual speech recognition systems often degrade in real-world scenarios due to signal corruption and distribution shifts. To address this, we propose a unified uncertainty-modeling framework, namely the uncertainty-aware Bayesian gating network (UBG-Net). UBG-Net features a Modality Uncertainty-aware Bayesian Fusion (MUBF) mechanism that injects signal-level aleatoric uncertainty into a Bayesian network to model epistemic uncertainty, thereby ensuring robust fusion of pre-trained backbone features. For inference, we introduce Distribution Uncertainty-aware Hierarchical Voting (DUHV) to select transcripts from Monte Carlo samples, prioritizing frequency and using inference scores in case of a tie. Experiments on the AVCocktail and LRS2 datasets demonstrate the overall superiority of UBG-Net compared to SOTA baselines. Ablation studies confirm that MUBF and DUHV effectively filter noise, enhancing fusion and decoding robustness.

View source

Similar papers

Open access Aug 2026

AV-DeepFake-Net: Attention-Guided and Uncertainty-Aware Network for Audiovisual DeepFake Detection

Detecting audio-visual DeepFake (AVDeepFake) is becoming increasingly important as synthetic media tools become widely accessible and spread across consumer devices. In this study, we present an Detecting audio-visual DeepFakes (AV-DeepFakes) has become increasingly critical with the rapid proliferation of accessible s...

Nasir Saleem, Ahmad Ali, Zhuo-Qi Zeng et al. · 0 citations
Preprint Sep 2026

A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration

Speaker verification back-ends commonly combine similarity scoring, score normalization, and calibration. However, speaker embeddings extracted from real-world utterances have trial-dependent reliability because of factors such as duration, noise, and channel variation. Existing uncertainty-aware methods primarily impr...

Junjie Li, K. Lee · 0 citations
Open access 2026

Location-Aware Mixture-of-Experts Framework for Adaptive Speech Enhancement

Speech enhancement models often struggle to adapt to heterogeneous and ambiguous noise conditions. This paper proposes a location-aware mixture-of-experts framework that jointly uses acoustic observations and categorical location context for adaptive speech enhancement. A gating network estimates continuous expert weig...

D. Noh, S. Kim · 0 citations
Preprint Aug 2026

Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition

Large language model (LLM)-based audio-visual speech recognition (AVSR) systems are robust under noise. Contrastive decoding (CD), originally introduced to stabilize LLM generation by contrasting a weaker model against a stronger one at inference time, adjusts predictions without additional training. In this work, we a...

Youngchae Kim, Dali Yang, Joon-Hyuk Chang · 0 citations
Conference Aug 2026

CACF-Net: A Cross-Attention Cognitive Fusion Network for Robust and Interpretable Signal Recognition

With the development of wireless communications, the electromagnetic environment has become increasingly complex. Reliable signal recognition is fundamental for spectrum awareness and intelligent management. However, single-domain features are vulnerable to noise and fading, leading to limited performance. To address t...

Xuan-Di Qiao, Han Zhang, Yi Jiang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.