Skip to content

Causal Supervision of Attention for Affective Behaviour Analysis

Jul 2026 · arXiv.org · Vol abs/2607.12091 · 0 citations · 48 references
Computer Science

TL;DR

An attention pooling framework that combines causal supervision with cross-covariance regularization of attention components is proposed, encouraging subject-invariant attention and non-redundant representations that improve generalization.

Abstract

The \textit{11th Affective Behaviour Analysis in-the-wild Competition} includes the Multi-Task Learning Challenge, where participants develop a unified framework for Valence-Arousal Estimation, Expression Recognition, and Action Unit Detection. The challenge lies in learning emotion-related representations that generalize across subjects while remaining robust to spurious factors such as identity, illumination, pose, and demographic variation. To aggregate features extracted by a pre-trained backbone into a compact representation for prediction, attention mechanisms selectively weight the most informative facial regions. However, these attention weights can still capture dataset-specific correlations rather than genuine affective cues. To address this limitation, we propose an attention pooling framework that combines causal supervision with cross-covariance regularization of attention components, encouraging subject-invariant attention and non-redundant representations that improve generalization. Our method achieves $CCC_{VA}=0.5123$ for VA estimation on the official validation set, together with $F_{EX}=0.3116$ and $F_{AU}=0.3974$ for expression recognition and action unit detection, respectively, resulting in an overall $P$ score (the sum of the individual task metrics) of $1.2214$.

View source

Similar papers

Jul 2026

AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis

This work presents a system for the Multi-Task Learning (MTL) track of the 11th Affective Behavior Analysis in-the-wild (ABAW) competition on s-Aff-Wild2, the static selected-frame version of Aff-Wild2, showing that post-encoder adaptation and task-wise modeling choices provide a strong MTL pipeline without training a...

Dipit Saha, Mohammad Raihan Rashid, Shahruz Mannan et al. · 0 citations
Open access Aug 2026

Emotion recognition from body movement through interpretable motion-aware sequential modeling

A Window Transformer architecture grounded in the Multiple Instance Learning (MIL) paradigm, which decomposes the full sequence as a single temporal stream into overlapping windows and learns to assign greater relevance to those segments containing stronger emotional content, providing a more interpretable framework.

Sergio Esteban-Romero, Iván Martín-Fernández, R. San-Segundo et al. · 0 citations
Book Open access Aug 2026

Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning

A multimodal framework that integrates the implicit spatiotemporal features extracted from pretrained video, audio, and image encoders along with structured behavioral modalities like head pose, gaze, facial action units, emotion, and wavelet-based audio features, demonstrating that reliable risk-quantification is an e...

Alperen Kantarcı, Visvanathan Ramesh, Gemma Roig · 1 citation
Preprint Aug 2026

LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation

LG-GER is proposed, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence that achieves competitive or superior results compared to state-of-the-art methods that require detection and multi-stream processing at inference.

Ahmed-Shehab Khan, Zhiyuan Li, Yan Tong · 0 citations
Open access 2026

MM-PSYCHE: Multimodal Multitask Psychological Characteristic Estimation Through Cross-Domain Semi-Supervised Learning

Psychological characteristic estimation from multimodal in-the-wild behavior is usually studied using separate corpora, each annotated for a single target task. Such annotation fragmentation limits cross-task learning and cross-domain generalization across affective, dispositional, and interactional phenomena. To addre...

E. Ryumina, A. Axyonov, D. Koryakovskaya et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.