Dual-Stream Decoupling of Sentiment and Facts for Multi-Modal Forgery Detection
Abstract
With the rapid development of Generative Artificial Intelligence (GAI) technology, multimodal fake media has spread widely on the Internet, posing a serious threat to the social information ecosystem. Existing multimodal manipulation detection methods mostly adopt a holistic feature fusion strategy, which is difficult to effectively capture fine-grained emotional conflicts such as Text Attribute Manipulation. To address this problem, this paper proposes a Dual-Stream Decoupled Multimodal Reasoning Network under the framework of the Detecting and Grounding Multi-Modal Media Manipulation task. This network decomposes multimodal information into an objective fact stream and a subjective emotion stream for separate processing: on the text side, fact features and emotion features are separated through a gating mechanism; on the visual side, contextual action features and facial micro-expression features are extracted correspondingly. On this basis, a directed cross-attention mechanism is designed to achieve accurate matching between fact features and contextual features, as well as between emotion features and micro-expression features. Experimental results on the Dual-Stream Decoupled Multimodal dataset show that the proposed method outperforms the original Hierarchical Multi-modal Manipulation Reasoning Transformer model on all core indicators: the overall F1-score of multi-label classification is improved by 0.63%, the F1-score of text token localization is improved by 0.97%, and in particular, the F1-score of the most subtle Text Attribute Manipulation is improved by 2.02%. This work effectively alleviates the performance defects of the original HAMMER model on fine-grained emotional manipulation detection, and provides a lightweight compatible improvement scheme targeting covert sentiment-based Text Attribute Manipulation. It achieves obvious metric gains on the hardest-to-detect TA category while maintaining comprehensive performance across all subtasks.