Skip to content
#small language model Open access

MiLTL: A Cross-Modal Contradiction Cascade for On-Device Voice Phishing Detection

Unknown authors
Aug 2026 · Algorithms · 0 citations

TL;DR

This work constructs KorMMP, a benchmark keeping real regulator-sourced scam audio and real benign speech but decorrelating transcript from label, and MiLTL, an on-device detector built on affective and neutrosophic channels, including a cross-modal contradiction signal formulated to weigh lexical warmth against vocal coldness.

Abstract

Voice phishing (vishing) causes financial and psychological harm, yet deployed detectors that scan transcripts for scam vocabulary are more fragile than their benchmark scores suggest: their near-perfect accuracy on standard Korean corpora is memorization, collapsing under scammer paraphrase or language-model rewriting. We construct KorMMP, a benchmark keeping real regulator-sourced scam audio and real benign speech but decorrelating transcript from label, and MiLTL, an on-device detector built on affective and neutrosophic channels, including a cross-modal contradiction signal (XM) formulated to weigh lexical warmth against vocal coldness. In the reported evaluation, XM’s measured contribution is confidence banding and escalation routing rather than ranking accuracy. A scoring rule with zero gradient-learned parameters at inference screens every segment, referring ambiguous calls to a small on-device language model; raw audio stays on the device. We measure both stages on a commodity CPU container and smartphone; these budgets cover the detector only and exclude speech recognition, which we expect to dominate a live end-to-end budget. With no training on the hard benchmark, MiLTL achieves an AUROC of 0.965 (recording-level cluster bootstrap 95% CI 0.943–0.984), while every evaluated comparator stays below 0.69: text-only, audio-only, fusion, 7B multimodal. MiLTL scores lower on the saturated corpus, as expected for a detector designed not to rely primarily on lexical shortcuts. KorMMP is a controlled stress test of lexical decorrelation, not a measure of in-the-wild detection; its harmful and benign audio come from different corpora, so the source and label are structurally confounded, and the residual source effects cannot be fully excluded. XM is author-defined, constructed rather than naturally occurring in the synthetic stratum and not yet validated against human perception; this is the principal open limitation of the work. Within that scope, the work provides a reproducible basis for on-device vishing defense and a benchmark for detector robustness under vocabulary shift.

Read PDF

Similar papers

Open access Sep 2026

VISH-GUARD: a multi-agent and LLM-powered framework for multilingual voice phishing detection

Voice phishing (vishing) has emerged as a major cybersecurity threat, leveraging persuasive speech and psychological manipulation to deceive victims in real time. Existing detection approaches remain limited by monolingual assumptions, unimodal analysis, and poor interpretability, reducing their effectiveness against...

Yasser Hmimou, Mohamed Tabaa, Azeddine Khiat et al. · 0 citations
Oct 2026

Evasion-hate: targeted augmentation for hate speech detection under prompt-induced surface-form shifts

Online hate speech detection is increasingly used as a decision-support component in content moderation, but most detectors are evaluated on clean held-out text. In practice, users can preserve harmful intent while changing surface form through spelling obfuscation, coded wording or code-switching, creating a robus...

Thanh Le, Thien Khai Tran · 0 citations
Conference Open access Sep 2026

ICFD-31k: A Large-Scale Dataset and Benchmark for Real-Time Conversational Fraud Detection

The proliferation of sophisticated telephone scams poses a significant societal and economic threat, impacting diverse linguistic contexts in a country like India. Furthermore, the lack of large-scale, publicly available datasets remains a critical barrier impacting research on robust, real-time countermeasures. In vie...

Rishi Ahuja, K. Prateek, Simranjit Singh · 1 citation
#natural language process... Preprint Sep 2026

CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection

Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation Harness), a bilingual counterfactual diagnostic framework that evaluates each utterance under orig...

Qi-Yang Sun, Xu-Dong Li, Yu-Pei Li et al. · 1 citation
Conference Open access 2026

LLM-Augmented Machine Learning for Phishing Email Detection: Enhancing Classification, Explainability, and Multilingual Support

Nowadays, email is the world’s primary communication channel, making it a dominant vector for phishing attacks that exploit urgency, authority, and trust. Advances in Large Language Models (LLMs) have intensified this threat by enabling attackers to generate persuasive, context-aware, and multilingual phishing emails a...

Farouk Aziz, Norbert Oláh · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.