Self-Supervised Representation Learning for Heterogeneous Behavioral Expressions of Social Engagement
Abstract
Social engagement is a complex construct that includes multiple behavioral signals. These signals vary across contexts, age groups, and clinical populations. Because of this variability, treating engagement as a single binary state, engaged versus not engaged, is difficult and often inconsistent. In this work, we introduce a self-supervised framework that learns visual representations of interaction dynamics directly from paired videos of two people interacting. Rather than predicting engagement labels directly, the model learns a core representation of how participants behave with and respond to each other. This shared representation can then be adapted, through fine-tuning, to predict multiple observable behaviors associated with engagement. Our model uses a two-stream video encoder, with one stream for each interaction partner. It is pretrained using a simple self-supervised task: distinguishing silence from active listening. This task requires no manual annotations and encourages the model to capture meaningful inter-participant dynamics. Without changing the model architecture, we transfer the pretrained encoder to three downstream tasks across three different datasets: focus of visual attention detection, backchannel detection, and perceived engagement estimation. Across all tasks, the pretrained two-stream encoder consistently outperforms both a single-stream baseline and a randomly initialized identical two-stream encoder. For perceived engagement estimation, our visual-only model achieves a concordance correlation coefficient (CCC) of 0.745, which is comparable to state-of-the-art multimodal audiovisual approaches. For focus of visual attention detection, the model reaches 80.24% accuracy. On the highly challenging task of backchannel detection, the pretrained encoder achieves 57.98% balanced accuracy, despite never being trained with backchannel-specific annotations. Ablation experiments show that both explicit modeling of inter-participant behavior and self-supervised pretraining are critical to performance. Together, these results suggest that diverse behavioral markers of engagement share a common foundation in interaction dynamics, and that this foundation can be learned without engagement-specific supervision. Project page and code will be available at https://compsygroup.github.io/ssl.social.engagement/.