Skip to content
Open access

TurkPhish v2: A Template-Composed Generation Matrix for Turkish Phishing E-Mail Corpora, with a Leakage Audit and a Detection Benchmark

Aug 2026 · Applied Sciences · Vol 16, pp. 8274 · 0 citations · 18 references

TL;DR

This work introduces TurkPhish v2, a class-balanced corpus of 2653 Turkish e-mails built from a generation matrix crossing phishing themes, persuasion tactics, length bands, and stylistic registers; the matrix is encoded as LLM-ready prompts, while the released corpus is instantiated by a deterministic, anti-leakage template composer.

Abstract

Phishing remains the dominant initial-access vector in modern cyberattacks, yet Turkish-language resources for benchmarking phishing e-mail detectors are scarce and methodologically fragile. We first show that a publicly released Turkish phishing dataset—our own earlier one (v1; 7504 messages)—is pathological: after entity masking it collapses to seven phishing and one legitimate content templates, and 99.95% of sender addresses share one artificial pattern, so any classifier memorizes artifacts rather than phishing semantics. We introduce TurkPhish v2, a class-balanced corpus of 2653 Turkish e-mails built from a generation matrix crossing phishing themes, persuasion tactics, length bands, and stylistic registers; the matrix is encoded as LLM-ready prompts, while the released corpus is instantiated by a deterministic, anti-leakage template composer. Against the v1 pathologies, the corpus is clean: unique masked templates, no near-duplicates (maximum pairwise cosine 0.808), opening and punctuation views at chance (49.8–52.8%), and matched lengths (p = 0.999). We also report the residual these gates miss: class signal lives in a closed pool of intent sentences, so a zero-learning lookup rule reaches 0.928 macro F1 in distribution and 0.894 out of topic, outscoring five of the ten benchmarked detectors. In-distribution performance is therefore a ceiling artifact, not evidence of learned phishing semantics. Benchmarking ten detectors across classic machine learning, Turkish/multilingual transformers, and an instruction-tuned large language model, eight exceed 0.99 macro F1 in distribution and are statistically indistinguishable (McNemar, p > 0.05), whereas out of topic TF-IDF linear models lead (logistic regression 0.947, 95% CI [0.926, 0.966]; 0.959 under the body-only protocol we recommend to users of the corpus) and mBERT collapses (0.570). Because persuasion phrasing is shared across themes, this protocol measures transfer to unseen theme vocabulary rather than robustness to novel phrasing. The transformer deficit out of topic is, for two of three encoders, a thresholding rather than a ranking failure: mBERT retains 0.960 out-of-topic AUC while its recall at the default 0.5 rule falls to 0.263. The corpus, splits, generation framework, and the full leakage audit, including its negative results, are available for research use.

Read PDF

Similar papers

Open access Aug 2026

Phishing GAT: Adversarial-Hardened Phishing Email Detection via Semantic-Structural Fusion and Graph Attention Networks

PhishingGAT, a detector that fuses word-level semantic features with structural ones and is hardened against adversarial perturbation, is presented, a detector that fuses word-level semantic features with structural ones and is hardened against adversarial perturbation.

R. Kodali, Siva Rama Krishna T Dr · 0 citations
Review Open access Aug 2026

A Multi-Vector Framework for Localized Phishing Detection URLs: Integrating Telegram-Sourced Intelligence and Iraqi Contextual Features

The outcome indicates that phishing detection improves when models incorporate attributes that match the target region, and the benefit is largest in emerging digital markets where attackers combine local social engineering with rapid infrastructure changes.

Jaber M. Al-Dulimi · 0 citations
Review Open access Sep 2026

A Real CVE-Backed Benchmark for WordPress Plugin Vulnerability Detection: Re-Evaluating Static and Learning-Based Detectors Under Leakage-Controlled Evaluation

Background: WordPress plugins account for the large majority of disclosed CMS vulnerabilities, and learning-based detectors report high accuracy on synthetic corpora and random splits. Methods: We build a benchmark from 1757 real plugin CVEs (6666 indexed identifiers, 1281 plugins), yielding 30,860 labeled PHP function...

Z. Tashenova, Aisultan Aitmagambetuly, A. Urynbassarova et al. · 0 citations
Review Open access Aug 2026

TrustPhish-AI: Safety-Gated Source-Aware Evaluation of Three-Class Phishing Detection Under Email-to-SMS Shift

Source shift can make phishing detectors appear stronger while concealing unsafe class-specific errors. This study evaluated a three-class workflow (Legitimate, Suspicious, and Phishing) using source-aware governance, controlled augmentation, family-disjoint diagnostics, and pre-specified safety gates. The frozen train...

Akam Aziz, Umran Abdullah Haje · 0 citations
Preprint Sep 2026

A Large-Scale Empirical Study of Modern Phishing Email Content

Phishing remains one of the most pervasive threats to Internet users, and email remains its predominant delivery channel. Email content is the attack surface of phishing: it is what the victim reads and what automated defenses inspect. Yet the composition of modern phishing content is poorly measured. Prior work has ch...

Jaehwan Park, Woonghee Lee, Fu-Jiao Ji et al. · 0 citations
Open access Sep 2026

Comparative Analysis of Transformer-Based and ClassicalMachine Learning Models for Phishing Email Detection:A Multi-Source Dataset Evaluation with Explainability

This paper compares five classifiers: Multinomial Naive Bayes, Random Forest, Bidirectional Long Short-Term Memory, BiLSTM, DistilBERT, and BERT-base, and finds that BERT-base achieves the highest F1-score and DistilBERT the lowest, representing the strongest accuracy–latency trade-off in the evaluated environment.

Andre Sebastian Samaniego Buñay, Ariel Misael Orellana Albarracin, Joel Marcelo Chuquimarca Pomagualli · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.