Skip to content
Review Open access

Audio-Visual-Textual Fusion for Emotion Recognition: A Neurologically Inspired Attention Mechanism

Aug 2026 · Eduschool Journal of Artificial Intelligence Research (EJAIR) · 0 citations · 16 references

Abstract

Systems that recognise emotion from speech, facial behaviour and language combine the three streams with mechanisms borrowed from machine translation rather than from any account of how the brain performs the same task. This article sets out a fusion mechanism derived from four established findings in multisensory neuroscience, namely the principle of inverse effectiveness, the temporal binding window, precision weighted cue combination and dual route affective processing. Each finding is translated into a computational operator, and the operators are assembled into an attention mechanism in which the weight given to cross-modal evidence rises as the reliability of a single stream falls, integration is restricted to a learned temporal window, contributions are scaled by estimated precision, and a fast coarse pathway supplies a prior that a slower pathway refines. We review the fusion families in current use, place published methods against the neuroscientific principles they implicitly satisfy, and specify the architecture in enough detail for direct implementation. The evaluation protocol needed to test the proposal is described, including the reliability degradation and modality removal conditions under which the predicted advantages should appear. The mechanism is presented as a design contribution supported by argument and by published evidence, and empirical validation on standard corpora remains to be carried out.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.