An attention pooling framework that combines causal supervision with cross-covariance regularization of attention components is proposed, encouraging subject-invariant attention and non-redundant representations that improve generalization.
Abstract
The \textit{11th Affective Behaviour Analysis in-the-wild Competition} includes the Multi-Task Learning Challenge, where participants develop a unified framework for Valence-Arousal Estimation, Expression Recognition, and Action Unit Detection. The challenge lies in learning emotion-related representations that generalize across subjects while remaining robust to spurious factors such as identity, illumination, pose, and demographic variation. To aggregate features extracted by a pre-trained backbone into a compact representation for prediction, attention mechanisms selectively weight the most informative facial regions. However, these attention weights can still capture dataset-specific correlations rather than genuine affective cues. To address this limitation, we propose an attention pooling framework that combines causal supervision with cross-covariance regularization of attention components, encouraging subject-invariant attention and non-redundant representations that improve generalization. Our method achieves $CCC_{VA}=0.5123$ for VA estimation on the official validation set, together with $F_{EX}=0.3116$ and $F_{AU}=0.3974$ for expression recognition and action unit detection, respectively, resulting in an overall $P$ score (the sum of the individual task metrics) of $1.2214$.
This work presents a system for the Multi-Task Learning (MTL) track of the 11th Affective Behavior Analysis in-the-wild (ABAW) competition on s-Aff-Wild2, the static selected-frame version of Aff-Wild2, showing that post-encoder adaptation and task-wise modeling choices provide a strong MTL pipeline without training a...
Dipit Saha, Mohammad Raihan Rashid, Shahruz Mannan et al.· arXiv.org· 0 citations
A Window Transformer architecture grounded in the Multiple Instance Learning (MIL) paradigm, which decomposes the full sequence as a single temporal stream into overlapping windows and learns to assign greater relevance to those segments containing stronger emotional content, providing a more interpretable framework.
Sergio Esteban-Romero, Iván Martín-Fernández, R. San-Segundo et al.· Frontiers in Artificial Inte...· 0 citations
A multimodal framework that integrates the implicit spatiotemporal features extracted from pretrained video, audio, and image encoders along with structured behavioral modalities like head pose, gaze, facial action units, emotion, and wavelet-based audio features, demonstrating that reliable risk-quantification is an e...
Alperen Kantarcı, Visvanathan Ramesh, Gemma Roig· Proceedings of the 28th Inte...· 1 citation
LG-GER is proposed, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence that achieves competitive or superior results compared to state-of-the-art methods that require detection and multi-stream processing at inference.
Ahmed-Shehab Khan, Zhiyuan Li, Yan Tong· 0 citations
Psychological characteristic estimation from multimodal in-the-wild behavior is usually studied using separate corpora, each annotated for a single target task. Such annotation fragmentation limits cross-task learning and cross-domain generalization across affective, dispositional, and interactional phenomena. To addre...
E. Ryumina, A. Axyonov, D. Koryakovskaya et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.