Skip to content

AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis

Jul 2026 · arXiv.org · Vol abs/2607.16546 · 0 citations · 41 references
Computer Science

TL;DR

This work presents a system for the Multi-Task Learning (MTL) track of the 11th Affective Behavior Analysis in-the-wild (ABAW) competition on s-Aff-Wild2, the static selected-frame version of Aff-Wild2, showing that post-encoder adaptation and task-wise modeling choices provide a strong MTL pipeline without training a new large-scale face foundation model.

Abstract

Affective behavior recognition in the wild requires joint prediction of continuous valence-arousal, categorical facial expression, and multi-label action units from unconstrained face images. We present our system for the Multi-Task Learning (MTL) track of the 11th Affective Behavior Analysis in-the-wild (ABAW) competition on s-Aff-Wild2, the static selected-frame version of Aff-Wild2. The method focuses on post-encoder adaptation: frozen AffectNet-supervised backbones provide multi-resolution features, while task-specific temporal heads and cross-task fusion modules select the useful signals for each target. For action-unit recognition, we adapt MAE-Face with Low-Rank Adaptation (LoRA) and use DISFA through per-unit expert routing rather than direct sequential transfer. Ablations over backbone, temporal, fusion, and AU-adaptation choices define the final configuration. The final system obtains P = 1.7302 on the official validation split, showing that post-encoder adaptation and task-wise modeling choices provide a strong MTL pipeline without training a new large-scale face foundation model.

View source

Similar papers

#natural language process... Preprint Sep 2026

Emotion as a Distribution: Joint Valence-Arousal Probability Learning for Speaker-Independent Multimodal Emotion Recognition

Human emotion is graded and frequently mixed, yet most multimodal recognizers collapse it onto a single hard label. We argue the recognizer should instead expose a distribution over affective space. Our text+speech system, alongside its categorical decision, emits a $9\times9$ probability matrix over the Valence-Arousa...

Ting-Yi Lin, Wen-Ren Yang, Kuan-Wei Chen · 0 citations
Open access Aug 2026

Toward Sustainable Cross-Time Affective Brain–Computer Interfaces: Explicit Modeling and Structural Calibration of Temporal Variability

Temporal variability limits the practical application of cross-time affective brain-computer interfaces (aBCI). Existing static feature-to-emotion mapping methods are unable to capture the dynamic changes in brain state. Here we propose a multi-level dynamic integrated perception network algorithm (MDIN), which expli...

Feifan Yan, Bo Zhang, Tao Wang et al. · 0 citations
Open access 2026

MM-PSYCHE: Multimodal Multitask Psychological Characteristic Estimation Through Cross-Domain Semi-Supervised Learning

Psychological characteristic estimation from multimodal in-the-wild behavior is usually studied using separate corpora, each annotated for a single target task. Such annotation fragmentation limits cross-task learning and cross-domain generalization across affective, dispositional, and interactional phenomena. To addre...

E. Ryumina, A. Axyonov, D. Koryakovskaya et al. · 0 citations
Open access Jul 2026

Emotion-BIND for multimodal emotion recognition and reasoning.

Constructing Multimodal Emotion Recognition in Conversation (MERC) models is important for understanding affective states from text, audio, images, and video. Existing approaches often rely on linear layers for cross-modal feature alignment and freeze encoder parameters during training, which can introduce feature degr...

Qunli Tang, Qi Wang, Shouhao Zhang et al. · 0 citations
Preprint Aug 2026

A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition

Mixed emotions represent a clinically relevant but still underexplored target for automatic emotion recognition. EEG provides millisecond-level access to neural activity, yet most EEG pipelines analyze the signal through a single temporal window, thereby fixing the temporal structure available to the model. This study...

Stefanos Gkikas, Yanglu Guo, Guangliang Li et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.