Skip to content
Open access

DAMSE: a dialect-aware multi-strategy ensemble framework for Arabic vishing detection with zero-shot learning

Jul 2026 · PeerJ Computer Science · Vol 12, pp. e3985 · 0 citations · 41 references
Computer Science

TL;DR

DAMSE is validated on two linguistically distinct datasets and establishes the first zero-shot baselines for vishing detection, introducing dialect-adaptive ensemble weighting that provides a consistent gain over uniform fusion, and releasing the first comprehensive multi-dialect Arabic vishing dataset with structured risk indicator annotations.

Abstract

Voice phishing (vishing) attacks targeting Arabic and Korean speakers represent a growing cybersecurity threat that demands automated detection systems capable of operating across diverse linguistic contexts. This article proposes Dialect-Aware Multi-Strategy Ensemble (DAMSE), a novel framework that synergistically fuses five complementary classification strategies for multilingual vishing detection: fine-tuned transformer models (Arabic BERT/Korean BERT), enhanced zero-shot classification with multi-template hypothesis averaging and Platt scaling calibration, few-shot learning via SetFit contrastive training, risk-score gradient boosting incorporating domain-specific social engineering indicators, and dialect-adaptive weighted ensemble fusion optimized per regional variant. We validate DAMSE on two linguistically distinct datasets: a curated Arabic vishing corpus comprising 448 multi-turn conversations across nine dialects (Modern Standard Arabic (MSA), Egyptian, Gulf, Jordanian, Saudi, Yemeni, Sudanese, Iraqi, Syrian) and 23 fine-grained categories, and the Korean KorCCVi v2 dataset containing 23,550 real-world call transcripts with severe class imbalance (1:13.3 ratio). Under stratified 5-fold cross-validation with repeated runs (five seeds × five folds = 25 runs), DAMSE achieves 99.31 ± 0.53% accuracy on Arabic and 99.92 ± 0.05% on Korean, demonstrating robust cross-lingual generalizability with tight variance bounds. On the held-out test splits, DAMSE achieves 99.99% accuracy on Arabic and 99.80% on Korean with zero false positives on both datasets. Key contributions include: establishing the first zero-shot baselines for vishing detection (92.00% Arabic, 94.20% Korean), achieving 96.50% accuracy with only five examples per class through data-efficient few-shot learning, introducing dialect-adaptive ensemble weighting that provides a consistent gain over uniform fusion, and releasing the first comprehensive multi-dialect Arabic vishing dataset with structured risk indicator annotations. Explainable artificial intelligence (AI) analysis reveals convergent vishing patterns across languages: financial terminology dominance, urgency markers, formal impersonation language, and extended narrative length. Comparison with 15 state-of-the-art approaches confirms DAMSE’s advancement across performance, zero-shot capability, multi-dialect support, and cross-lingual generalizability, establishing strong benchmarks for vishing detection research.

Read PDF

Similar papers

Review Open access Aug 2026

Enhancing dialectal Arabic aspect based sentiment analysis through a novel Algerian dialect telecommunication dataset using a unified end-to-end approach

A unified End-to-End (E2E) framework for Arabic ABSA based on the newly introduced dataset and the Arabic ABSA hotels dataset is proposed, where aspect term extraction and sentiment classification are integrated into a single sequence labeling task.

Fatiha Tebbani, C. Kara-Mohamed, A. Hamdi-Cherif · 0 citations
Preprint Aug 2026

AraSSM: A bidirectional state-space encoder for Arabic masked language modeling

A bidirectional Mamba encoder pretrained via masked language modeling on a corpus combining Arabic Wikipedia and CulturaX text is introduced, trained end-to-end on four consumer-grade NVIDIA RTX 2080Ti GPUs (11GB) over approximately ten days.

A. A. Aliane, H. Aliane, N. Semmar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.