Skip to content
Open access

Context-aware multimodal reasoning for explainable Bengali fake news detection using vision-language models

Aug 2026 · Discover Networks · Vol 2 · 0 citations · 30 references

TL;DR

The proposed context-aware multimodal reasoning approach for explainable Bengali fake news detection surpasses conventional CNN, transformer, and unimodal baselines on several performance metrics and indicates how context-based multimodal reasoning can improve the efficiency and robustness of the model, along with making it more interpretable.

Abstract

In particular, the rapid dissemination of fake news using social media platforms has posed a severe threat to social stability and information reliability in language groups that lack sufficient resources. In addition, it has become difficult to automatically identify fake news in the case of Bengali digital media because fake news is mostly distributed using multimodal data like memes, screenshots, posters, photos with texts, etc. The effectiveness of the existing fake news detection methods is somewhat hindered by their inability to provide explainability and their focus mainly on either textual or visual data. This study proposes a context-aware multimodal reasoning approach for explainable Bengali fake news detection. The proposed model incorporates EasyOCR for Bengali textual information extraction, ResNet-50 and ViT for supplementing visual feature learning, and Qwen2-VL-2B-Instruct for multimodal semantic inference. The proposed approach can detect semantic contradictions and associations between textual assertions and image information through matching textual and visual data based on a context-sensitive fusion technique. This methodology generates interpretable explanations beyond the fake/real binary categorization to enhance user trust in automated outcomes. According to an experiment conducted using a multimodal Bengali fake news dataset, the proposed approach surpasses conventional CNN, transformer, and unimodal baselines on several performance metrics. Results indicate how context-based multimodal reasoning can improve the efficiency and robustness of the model, along with making it more interpretable. The proposed method serves as an encouraging route towards the identification of fake news in multiple low resource languages, as well as in combating misinformation in the Bengali digital ecosystem.

Read PDF

Similar papers

Preprint Sep 2026

Detecting and Explaining Fake News Short Videos with Multimodal Content and Real-World Evidence

NVKE-CEI is a unified system that integrates a news video keyframes extraction method (NVKE) and an FNVDE framework leveraging both content and evidence information (CEI), which outperforms state-of-the-art baselines while generating high-quality content-grounded explanations.

Yi-Feng Luo, Yu-Peng Li, Ming Tang et al. · 0 citations
Aug 2026

KGEMD: knowledge-guided enhanced multimodal detection for fake news

The results indicate that explicitly modeling semantic conflict as a discriminative feature effectively improves detection precision and generalization, providing a robust solution for factual verification in complex media environments.

Zi-Heng Wang, Jun-Fang Song, Shu-Yu Wang et al. · 0 citations
Review Open access Sep 2026

Feature Based Survey on Fake News Detection: Statistical and Semantic

The digital news portals and social media are rapidly expanding, which has significantly increased the spread of fake news, which affects public opinion, social harmony, and trust in information sources. Detection of fake news at an early stage is a critical research challenge. In recent years, researchers have applied...

Itika U. Lakkewar, R. Jugele · 0 citations
Book Open access Aug 2026

MAR: Metacognitive Agentic Reasoning for Multimodal Fake News Detection

The Metacognitive Agentic Reasoning for Multimodal Fake News Detection (MAR), a framework that integrates multi-agent metacognitive debate with external knowledge retrieval to enhance the accuracy, generalization, and interpretability of multimodal fake news detection.

Wen-Yu Chen, Hengbing Dong, Junhao Wa et al. · 0 citations
Open access Sep 2026

Deep Learning Methods for Multimodal Fake News Classification Combining Textual and Visual Information

The authors suggest a computationally efficient multimodal deep learning framework using Bidirectional Encoder Representations of Transformers (BERT) to extract textual features and convolutional neural networks to learn visual representations that offers a computational scaling alternative to attention-based models, w...

P. Jadhav, R. K. Shukla · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.