Skip to content
Book Open access

A Multimodal Cross-Attention Framework for Robot Error Detection from Bystander Reactions

Oct 2026 · 0 citations

Abstract

To detect human-robot interaction errors, existing approaches have primarily focused on behavioural responses from individuals directly involved in the interaction, whereas reactions from external observers have received comparatively little attention. Similarly, these approaches have been limited to interactions in controlled environments with little variability in lighting, camera angles, and background noise. Addressing these gaps, the ERR@HRI 3.0 challenge focuses on bystander error detection in unconstrained, real-world environments. Our primary submission is a time-aware multimodal fusion framework that combines facial and acoustic representations via bidirectional cross-attention. Two versions of this framework, with different encoding approaches, are evaluated against a bidirectional LSTM (BiLSTM) classifier with simple concatenation. Our results demonstrate the effectiveness of this cross-fusion approach for capturing errors and generalising to unseen real-world data. Our cross-fusion and BiLSTM models achieved the top-three positions on Track 1, outperforming the challenge baseline and other competing submissions. Feature importance analysis further revealed that smile-related action units and their temporal dynamics serve as primary discriminative cues for error detection. Our code is available at https://github.com/chardiwall/SHEF-HRI.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.