A multimodal rumor detection framework that integrates frequency-domain features with sequential dependency modeling and an adaptive gated fusion strategy dynamically balances the contributions of different modalities, thereby enhancing the discriminative capability of the fused representation.
Abstract
With the rapid growth of social media, multimodal content combining text and images has become a major medium for rumor dissemination. However, existing multimodal rumor detection methods primarily focus on semantic modeling in the spatial domain, making it difficult to simultaneously capture latent image manipulation traces and structural dependencies among image regions. Moreover, frequency-domain features are susceptible to noise during feature extraction, which limits the model’s ability to identify forged content effectively. To address these challenges, this paper proposes a multimodal rumor detection framework that integrates frequency-domain features with sequential dependency modeling. For textual representation, a pretrained BERT encoder is employed to extract contextual semantic features. For visual representation, a spatial–frequency dual-branch architecture is designed. Specifically, the spatial branch combines ResNet34 and a Bi-GRU network to model sequential dependencies among image regions, thereby capturing long-range contextual relationships. The frequency-domain branch employs the Discrete Cosine Transform (DCT) and a residual convolutional network to extract potential frequency-domain anomalies, while a gating mechanism is introduced to suppress noise and enhance the representation of forgery-related features. Furthermore, a cross-modal cross-attention mechanism is employed to achieve bidirectional semantic alignment and model semantic inconsistencies between textual and visual modalities. Finally, an adaptive gated fusion strategy dynamically balances the contributions of different modalities, thereby enhancing the discriminative capability of the fused representation. Experimental results on two benchmark datasets demonstrate that the proposed method consistently outperforms existing baseline methods. On the Weibo dataset, the proposed method achieves an Accuracy of 95.41 ± 0.48% and a Macro-F1 of 95.36 ± 0.50%, while on the Pheme dataset, it achieves an Accuracy of 87.63 ± 0.58% and a Macro-F1 of 85.26 ± 0.82%.
THGNN-MRD, a temporal heterogeneous graph neural network for multimodal rumor detection, is proposed, suggesting that modeling rumors as evolving multimodal social events provides a principled and effective solution for trustworthy rumor detection.
The authors suggest a computationally efficient multimodal deep learning framework using Bidirectional Encoder Representations of Transformers (BERT) to extract textual features and convolutional neural networks to learn visual representations that offers a computational scaling alternative to attention-based models, w...
P. Jadhav, R. K. Shukla· International Research Journ...· 0 citations
Multi-domain multimodal fake news detection has attracted increasing research attention because misinformation on social media often involves both textual and visual content across heterogeneous topical domains. However, most existing methods assume that textual and visual modalities are simultaneously available, which...
Tingjuan Deng, Xiaolong Xu· Journal of King Saud Univers...· 0 citations
With the rapid development of Generative Artificial Intelligence (GAI) technology, multimodal fake media has spread widely on the Internet, posing a serious threat to the social information ecosystem. Existing multimodal manipulation detection methods mostly adopt a holistic feature fusion strategy, which is difficult...
Yi Zhou, Hao-Yang Wan, Ke-Yu Wu et al.· 2026 12th International Conf...· 0 citations
Understanding crowd behavior in video requires simultaneously reasoning over temporal dynamics, spatial appearance, and semantic context, challenges that single-modality models address only partially. Current vision-language systems suffer from a representation bottleneck originating not from encoder capacities, but fr...
The proliferation of social media has created fertile ground for misinformation, a challenge further intensified by recent advances in generative artificial intelligence. Modern fake news increasingly takes the form of sophisticated multimodal campaigns, where synthetic images and stylistically manipulated text are joi...
Mao-Lin Wang, Ziting Mai, Zi-Chun Liu et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.