RAMamba-Net is proposed, a reliability-aware Mamba-based multimodal fusion network for AAD that effectively exploits complementary EEG-EOG information, yielding accuracy gains over unimodal baselines, and is robust to signal perturbation and parameter variation.
Abstract
Auditory attention decoding (AAD) identifies the attended speaker from physiological signals, supporting neuro-steered hearing devices and natural human-machine interaction. Electroencephalography (EEG) is the dominant modality for AAD but provides incomplete evidence in naturalistic audio-visual scenes, motivating EEG and electrooculography (EOG) fusion. Existing approaches remain limited by weak cross-modal interaction, inefficient temporal modeling, and low robustness to sample variations. To address the limitations, we propose RAMamba-Net, a reliability-aware Mamba-based multimodal fusion network for AAD. RAMamba-Net employs a Mamba-enhanced band-aware convolutional Transformer to capture band-specific EEG patterns and long-range temporal dynamics. A dual-branch temporal-spatial encoder models EOG temporal and inter-channel dependencies. Cross-modal attention enables explicit modality interaction. Then, a reliability-aware module is introduced to estimate sample-wise modality weights for feature and prediction consistency, thereby enhancing multimodal fusion. Experiments on two AAD benchmarks demonstrate that RAMamba-Net effectively exploits complementary EEG-EOG information, yielding accuracy gains of 5.76% over unimodal baselines, together with more robust decoding and discriminative representations. Further analyses show that explicit cross-modal interaction improves multimodal alignment, while the reliability-aware module suppresses unreliable modality evidence and is robust to signal perturbation and parameter variation.
A dual-stream time-frequency convolutional network with additive attention that improves sensitivity to and modeling of frequency variation patterns and significantly outperform state-of-the-art methods is proposed.
Xiao Lin, Sunjie Zhang, Shuai Huang· Sheng wu yi xue gong cheng x...· 0 citations
Emotion-aware human–computer interaction increasingly relies on continuous emotion recognition (CER) to track affective states over time. This paper investigates continuous valence prediction on the MAHNOB-HCI database using a multimodal EEG+audio framework. The proposed model combines (i) a hierarchical spatiotemporal...
Ali Amini, Sarmad Maqsood, Irfan Abbas et al.· 0 citations
A lightweight and generalizable decoding framework named Hierarchical Convolutional Fusion Transformer (HCFT), which combines dual-branch convolutional encoders and hierarchical Transformer blocks for multi-scale EEG representation learning, and exhibits strong cross-subject generalization and structural interpretabili...
Haodong Zhang, Jiapeng Zhu, Yitong Chen et al.· IEEE journal of biomedical a...· 0 citations
Driver drowsiness is a major contributor to road accidents worldwide. Electroencephalography (EEG) enables direct measurement of neural correlates of cognitive fatigue, allowing early detection before behavioral symptoms manifest. This paper proposes a lightweight hybrid temporal–spectral EEG model with gated adaptive...
Arya S. Prakash, R. Resmi· 2026 Control Instrumentation...· 0 citations
Following a target speaker in a noisy environment, commonly known as the cocktail party problem, remains particularly challenging for cochlear implant (CI) users. Recent studies have explored EEG-based auditory attention decoding (AAD) using neural networks to enhance hearing assistance. This paper presents a resource-...
Qier Ma, R. George, Stefan Scholze et al.· 0 citations
EEG-based imagined speech classification is an important topic in brain–computer interface research. However, Chinese imagined speech EEG sig-nals are typically characterized by low signal-to-noise ratio, strong non-stationarity, and subtle inter-class differences, which make stable modeling challenging. Existing con...
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.