Skip to content
Open access

Linguistic markers of deception in Arabic news headlines: A cross-corpus study of stylistic and numeric features

Aug 2026 · PLoS ONE · Vol 21 · 0 citations · 40 references
Medicine

TL;DR

This study investigates automatic fake-news detection using Arabic headlines, leveraging five heterogeneous corpora and their unified combination, and proposes a late-fusion strategy coupling transformer representations with discriminative engineered features.

Abstract

The rapid spread of misinformation in Arabic news headlines poses a growing challenge to digital media integrity, given the scarcity of Arabic-specific automated detection tools relative to English-centric systems. Headlines are brief and context-limited, yet their lexical and stylistic patterns encode strong cues of veracity or deception, making headline-only detection practically urgent and linguistically tractable. This study investigates automatic fake-news detection using Arabic headlines, leveraging five heterogeneous corpora and their unified combination. English sets were incorporated via neural machine translation with light post-normalization that preserves stylistic cues, yielding a heterogeneous cross-domain corpus. A systematic analysis of linguistic and stylistic indicators reveals stable asymmetries between fake and real headlines that recur across domains. We evaluate approaches from classical TF-IDF baselines to Arabic-specialized transformers, and propose a late-fusion strategy coupling transformer representations with discriminative engineered features. Transformers consistently outperform classical baselines, confirming that subword representations effectively capture semantic and stylistic regularities in short Arabic texts. Late fusion yields statistically significant improvements only on datasets with prominent numeric or temporal cues; on the unified corpus, McNemar’s exact test confirms that fusion gains are non-significant, indicating that subword encoders already internalize the surface-level cues captured by the engineered features. Even where accuracy differences are marginal, interpretable features enhance explainability.

Read PDF

Similar papers

Open access Aug 2026

Unmasking Fake News: Lexical Patterns in COVID-19 Misinformation

The surge of fake news on social media during the COVID-19 pandemic posed serious risks to public health and social stability. While previous studies have largely focused on computational detection of misinformation, less attention has been paid to identifying human-recognisable linguistic features that enable real-tim...

Sharmin Haque, Adlina Ariffin · 0 citations
#natural language process... Preprint Sep 2026

CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection

Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation Harness), a bilingual counterfactual diagnostic framework that evaluates each utterance under orig...

Qi-Yang Sun, Xu-Dong Li, Yu-Pei Li et al. · 1 citation
Open access Sep 2026

Emotion recognition in cross-linguistic legal context: a comparison between human-based and computational approaches

Emotion shapes credibility assessments, judicial decision-making, and perceptions of procedural justice, yet its reliable detection in courtroom settings remains a significant methodological challenge. This study provides a multimethod evaluation of emotion-recognition approaches in cross-linguistic legal discourse usi...

Jing-Yi Li, Di-Fu Shi, Huolingxiao Kuang et al. · 0 citations
Open access Sep 2026

The Corpus of English–Tagalog Code-switching: an integrative corpus of linguistic contact effects

The Corpus of English–Tagalog Code-switching (CEnTaCS) is a new corpus designed to document naturalistic English–Tagalog bilingual speech. Locally referred to as Taglish, it is a widely-used yet understudied contact variety in the Philippines. Recorded in 2025 in Metro Manila, the corpus comprises three interrelated...

Aaron Santa Maria, Renata Enghels · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.