Skip to content
Open access

Fusion of Frequency-Domain Features and Sequential Dependency for Multimodal Rumor Detection

Aug 2026 · Algorithms · 0 citations · 38 references

TL;DR

A multimodal rumor detection framework that integrates frequency-domain features with sequential dependency modeling and an adaptive gated fusion strategy dynamically balances the contributions of different modalities, thereby enhancing the discriminative capability of the fused representation.

Abstract

With the rapid growth of social media, multimodal content combining text and images has become a major medium for rumor dissemination. However, existing multimodal rumor detection methods primarily focus on semantic modeling in the spatial domain, making it difficult to simultaneously capture latent image manipulation traces and structural dependencies among image regions. Moreover, frequency-domain features are susceptible to noise during feature extraction, which limits the model’s ability to identify forged content effectively. To address these challenges, this paper proposes a multimodal rumor detection framework that integrates frequency-domain features with sequential dependency modeling. For textual representation, a pretrained BERT encoder is employed to extract contextual semantic features. For visual representation, a spatial–frequency dual-branch architecture is designed. Specifically, the spatial branch combines ResNet34 and a Bi-GRU network to model sequential dependencies among image regions, thereby capturing long-range contextual relationships. The frequency-domain branch employs the Discrete Cosine Transform (DCT) and a residual convolutional network to extract potential frequency-domain anomalies, while a gating mechanism is introduced to suppress noise and enhance the representation of forgery-related features. Furthermore, a cross-modal cross-attention mechanism is employed to achieve bidirectional semantic alignment and model semantic inconsistencies between textual and visual modalities. Finally, an adaptive gated fusion strategy dynamically balances the contributions of different modalities, thereby enhancing the discriminative capability of the fused representation. Experimental results on two benchmark datasets demonstrate that the proposed method consistently outperforms existing baseline methods. On the Weibo dataset, the proposed method achieves an Accuracy of 95.41 ± 0.48% and a Macro-F1 of 95.36 ± 0.50%, while on the Pheme dataset, it achieves an Accuracy of 87.63 ± 0.58% and a Macro-F1 of 85.26 ± 0.82%.

Read PDF

Similar papers

#graph neural networks Open access Sep 2026

Temporal heterogeneous graph neural networks for multimodal rumor detection on social media

THGNN-MRD, a temporal heterogeneous graph neural network for multimodal rumor detection, is proposed, suggesting that modeling rumors as evolving multimodal social events provides a principled and effective solution for trustworthy rumor detection.

Dong Liu, Zhi-Yong Wang, Rui Zhang · 0 citations
Open access Sep 2026

Deep Learning Methods for Multimodal Fake News Classification Combining Textual and Visual Information

The authors suggest a computationally efficient multimodal deep learning framework using Bidirectional Encoder Representations of Transformers (BERT) to extract textual features and convolutional neural networks to learn visual representations that offers a computational scaling alternative to attention-based models, w...

P. Jadhav, R. K. Shukla · 0 citations
Open access Aug 2026

IMFND-AFR: an incomplete-modality fake news detection framework based on active feature reconstruction

Multi-domain multimodal fake news detection has attracted increasing research attention because misinformation on social media often involves both textual and visual content across heterogeneous topical domains. However, most existing methods assume that textual and visual modalities are simultaneously available, which...

Tingjuan Deng, Xiaolong Xu · 0 citations
Conference Aug 2026

Dual-Stream Decoupling of Sentiment and Facts for Multi-Modal Forgery Detection

With the rapid development of Generative Artificial Intelligence (GAI) technology, multimodal fake media has spread widely on the Internet, posing a serious threat to the social information ecosystem. Existing multimodal manipulation detection methods mostly adopt a holistic feature fusion strategy, which is difficult...

Yi Zhou, Hao-Yang Wan, Ke-Yu Wu et al. · 0 citations
Open access 2026

TIV: A Tri-Modal Text–Image–Video Architecture for Crowd Behavior Classification and Captioning

Understanding crowd behavior in video requires simultaneously reasoning over temporal dynamics, spatial appearance, and semantic context, challenges that single-modality models address only partially. Current vision-language systems suffer from a representation bottleneck originating not from encoder capacities, but fr...

Sheilla Wesonga, Jangsik Park · 0 citations
Book Open access Aug 2026

A Serial Two-Stage Framework for Robust Multimodal Fake News Detection via Adaptive Reasoning

The proliferation of social media has created fertile ground for misinformation, a challenge further intensified by recent advances in generative artificial intelligence. Modern fake news increasingly takes the form of sophisticated multimodal campaigns, where synthetic images and stylistically manipulated text are joi...

Mao-Lin Wang, Ziting Mai, Zi-Chun Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.