Emotion recognition from body movement through interpretable motion-aware sequential modeling
A Window Transformer architecture grounded in the Multiple Instance Learning (MIL) paradigm, which decomposes the full sequence as a single temporal stream into overlapping windows and learns to assign greater relevance to those segments containing stronger emotional content, providing a more interpretable framework.