Skip to content

RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

RAMamba-Net is proposed, a reliability-aware Mamba-based multimodal fusion network for AAD that effectively exploits complementary EEG-EOG information, yielding accuracy gains over unimodal baselines, and is robust to signal perturbation and parameter variation.

Abstract

Auditory attention decoding (AAD) identifies the attended speaker from physiological signals, supporting neuro-steered hearing devices and natural human-machine interaction. Electroencephalography (EEG) is the dominant modality for AAD but provides incomplete evidence in naturalistic audio-visual scenes, motivating EEG and electrooculography (EOG) fusion. Existing approaches remain limited by weak cross-modal interaction, inefficient temporal modeling, and low robustness to sample variations. To address the limitations, we propose RAMamba-Net, a reliability-aware Mamba-based multimodal fusion network for AAD. RAMamba-Net employs a Mamba-enhanced band-aware convolutional Transformer to capture band-specific EEG patterns and long-range temporal dynamics. A dual-branch temporal-spatial encoder models EOG temporal and inter-channel dependencies. Cross-modal attention enables explicit modality interaction. Then, a reliability-aware module is introduced to estimate sample-wise modality weights for feature and prediction consistency, thereby enhancing multimodal fusion. Experiments on two AAD benchmarks demonstrate that RAMamba-Net effectively exploits complementary EEG-EOG information, yielding accuracy gains of 5.76% over unimodal baselines, together with more robust decoding and discriminative representations. Further analyses show that explicit cross-modal interaction improves multimodal alignment, while the reliability-aware module suppresses unreliable modality evidence and is robust to signal perturbation and parameter variation.

View source

Similar papers

Book Open access Oct 2026

Multi-Scale Spatiotemporal EEG and Self-Supervised Audio Fusion: A Mixture-of-Experts Approach to Continuous Affect

Emotion-aware human–computer interaction increasingly relies on continuous emotion recognition (CER) to track affective states over time. This paper investigates continuous valence prediction on the MAHNOB-HCI database using a multimodal EEG+audio framework. The proposed model combines (i) a hierarchical spatiotemporal...

Ali Amini, Sarmad Maqsood, Irfan Abbas et al. · 0 citations
Aug 2026

HCFT: A Hierarchical Convolutional Fusion Transformer for Cross-Task EEG Decoding.

A lightweight and generalizable decoding framework named Hierarchical Convolutional Fusion Transformer (HCFT), which combines dual-branch convolutional encoders and hierarchical Transformer blocks for multi-scale EEG representation learning, and exhibits strong cross-subject generalization and structural interpretabili...

Haodong Zhang, Jiapeng Zhu, Yitong Chen et al. · 0 citations
Conference Aug 2026

A Lightweight Hybrid Temporal–Spectral Model with Gated Fusion for EEG-Based Driver Drowsiness Detection

Driver drowsiness is a major contributor to road accidents worldwide. Electroencephalography (EEG) enables direct measurement of neural correlates of cognitive fatigue, allowing early detection before behavioral symptoms manifest. This paper proposes a lightweight hybrid temporal–spectral EEG model with gated adaptive...

Arya S. Prakash, R. Resmi · 0 citations
Preprint Aug 2026

A Resource-Efficient CNN-Based EEG Auditory Attention Decoding ASIC

Following a target speaker in a noisy environment, commonly known as the cocktail party problem, remains particularly challenging for cochlear implant (CI) users. Recent studies have explored EEG-based auditory attention decoding (AAD) using neural networks to enhance hearing assistance. This paper presents a resource-...

Qier Ma, R. George, Stefan Scholze et al. · 0 citations
Conference 2026

Chinese Imagined Speech EEG Classification Method Based on Topological and Frequency-Band Priors

EEG-based imagined speech classification is an important topic in brain–computer interface research. However, Chinese imagined speech EEG sig-nals are typically characterized by low signal-to-noise ratio, strong non-stationarity, and subtle inter-class differences, which make stable modeling challenging. Existing con...

浩然 郭 · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.