Skip to content
Preprint

LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

LG-GER is proposed, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence that achieves competitive or superior results compared to state-of-the-art methods that require detection and multi-stream processing at inference.

Abstract

Inferring the collective emotional state of a group of people from a single image, a task known as group emotion recognition (GER), requires integrating spatially distributed cues such as faces, poses, interactions, and scene context. Current methods rely on detector-driven multi-stream pipelines. These are trained with only image-level supervision that lacks guidance on which regions matter or how strongly each contributes. We propose LG-GER, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence, i.e., bounding boxes paired with emotion signals and confidence scores, for the training images. This structured evidence is distilled into a single vision-language model (VLM) backbone through four complementary losses: classification, region-text grounding, spatial emotion, and spatial confidence regression. At inference, LG-GER requires no detectors, no MLLM, and no multi-stream fusion, making GER practical for real-time and resource-constrained deployment. LG-GER has been evaluated on two benchmark GER datasets (GroupEmoW and GAF~3.0) and achieves competitive or superior results compared to state-of-the-art methods that require detection and multi-stream processing at inference.

View source

Similar papers

Open access Aug 2026

Large Language Model-Assisted Distillation–Fusion Framework for Visual Emotion Recognition

A large language model-assisted distillation–fusion framework (VERLADF) is proposed, which introduces emotion instruction data generated by GPT to fine-tune a VLM, thereby enhancing its emotional semantic understanding capability and adaptively fuses predictions from the instruction-tuned VLM and the distillation modul...

Yujun Ma, Yun-Jie Zeng, Zhi-Yuan Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

BiCFlow-MER: Orchestrating Discriminative and Generative Multimodal Emotion Recognition via Conditional Transport

In multimodal emotion recognition (MER), human affective states are inferred by integrating complementary cues from multiple modalities. In audio-text MER, affective cues are often entangled with speaker style and lexical content, while cross-modal disagreement further complicates how the evidence should be integrated....

Yan-Bin Wang, Shen-Yue Wang, Chun-Yang Yu · 0 citations
Review Open access Aug 2026

Audio-Visual-Textual Fusion for Emotion Recognition: A Neurologically Inspired Attention Mechanism

Systems that recognise emotion from speech, facial behaviour and language combine the three streams with mechanisms borrowed from machine translation rather than from any account of how the brain performs the same task. This article sets out a fusion mechanism derived from four established findings in multisensory neur...

Navin Chandran · 0 citations
Preprint Sep 2026

Decoding Affective Nuances: Enhancing MLLMs via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning

While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap: MLLMs are difficult to reliably distinguish semantically p...

C. Ye, Wei-Dong Chen, Zhao-Bo Qi et al. · 0 citations
Conference Aug 2026

Multimodal Emotion Recognition with Emotion-Specific Cross-Modal Attention Blocks

Multimodal emotion recognition is increasingly important for healthcare, education, and human-computer interaction. However, many existing systems learn a single shared representation for all emotions, which can blur subtle class-specific cues. This paper proposes an emotion-specific multimodal architecture that combin...

Gnanaseelan Dharshika, A. Ramanan · 0 citations
#natural language process... Preprint Sep 2026

Emotion as a Distribution: Joint Valence-Arousal Probability Learning for Speaker-Independent Multimodal Emotion Recognition

This text+speech system, alongside its categorical decision, emits a $9\times9$ probability matrix over the Valence-Arousal plane, trained with a two-dimensional Gaussian soft target under a Kullback-Leibler/cross-entropy objective, aimed at counseling support.

Ting Lin, Wen-Ren Yang, Kuan-Wei Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.