What Matters Over Time: Temporal Properties of Non-Verbal Signals for Intent Recognition in Human–Robot Interaction
Abstract
Multimodal fusion systems for human–robot interaction predominantly rely on snapshot-level features that capture what nonverbal signals are present at a given moment, ignoring when they occur, how long they last, and in what order they unfold. Nevertheless, the temporal dynamics of human behaviour may convey communicative intent beyond the presence of individual cues. This paper investigates the temporal properties of nonverbal signals including onset timing, duration, ordering, and cross-modal coordination across gaze, gesture, and speech for intent recognition in human-robot interaction. We empirically characterise how these properties differ across 30 communicative intent classes in MIntRec 2.0, finding that speech timing features are the strongest discriminators, that speech act function governs nonverbal temporal patterns more than valence or arousal, and that cross-modal ordering follows a consistent intent-modulated hierarchy. Our work argues that the bottleneck in current multimodal fusion lies not in the architecture but in a feature space that discards the temporal structure most relevant to social understanding. Future work will incorporate these features into a multimodal fusion model and evaluate their effect on robot intent recognition in real HRI settings.