Skip to content
Open access

Multimodal Decoupling Chain of Thought for Affective Understanding

Sep 2026 · Electronics · 0 citations · 21 references

Abstract

The Chain-of-Thought (CoT) technique has significantly improved the affective reasoning capabilities of large language models (LLMs). However, existing multimodal CoT typically aggregate information from different modalities into mixed inputs, which allows the model to perform entangled reasoning. This paradigm often exhibits cross-modal information confusion and misalignment when dealing with complex sentiment samples with semantic conflicts. To address these challenges, we propose a Multimodal Decoupling Chain-of-Thought for affective understanding called MDCoT. Specifically, MDCoT structures LLM reasoning into three explicit and progressive steps: (1) Shared Sentiment Aggregation, which extracts common emotional baselines aligned across all modalities; (2) Private Sentiment Disentanglement, which isolates modality-specific emotional nuances, such as textual metaphors, facial micro-expressions, or audio tone shifts; and (3) Global Integrated Inference, which balances shared consensus and modal-private clues to make the final prediction. Extensive experiments across three affective tasks (i.e., sentiment analysis, sarcasm detection, intent recognition) demonstrate that MDCoT consistently outperforms state-of-the-art end-to-end models and existing prompting strategies across multiple MLLMs. Furthermore, ablation and robustness studies confirm that explicit disentanglement effectively prevents cross-modal conflict and improves decision stability.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.