Aug 2026· Applied Sciences· Vol 16, pp. 8274· 0 citations· 18 references
TL;DR
This work introduces TurkPhish v2, a class-balanced corpus of 2653 Turkish e-mails built from a generation matrix crossing phishing themes, persuasion tactics, length bands, and stylistic registers; the matrix is encoded as LLM-ready prompts, while the released corpus is instantiated by a deterministic, anti-leakage template composer.
Abstract
Phishing remains the dominant initial-access vector in modern cyberattacks, yet Turkish-language resources for benchmarking phishing e-mail detectors are scarce and methodologically fragile. We first show that a publicly released Turkish phishing dataset—our own earlier one (v1; 7504 messages)—is pathological: after entity masking it collapses to seven phishing and one legitimate content templates, and 99.95% of sender addresses share one artificial pattern, so any classifier memorizes artifacts rather than phishing semantics. We introduce TurkPhish v2, a class-balanced corpus of 2653 Turkish e-mails built from a generation matrix crossing phishing themes, persuasion tactics, length bands, and stylistic registers; the matrix is encoded as LLM-ready prompts, while the released corpus is instantiated by a deterministic, anti-leakage template composer. Against the v1 pathologies, the corpus is clean: unique masked templates, no near-duplicates (maximum pairwise cosine 0.808), opening and punctuation views at chance (49.8–52.8%), and matched lengths (p = 0.999). We also report the residual these gates miss: class signal lives in a closed pool of intent sentences, so a zero-learning lookup rule reaches 0.928 macro F1 in distribution and 0.894 out of topic, outscoring five of the ten benchmarked detectors. In-distribution performance is therefore a ceiling artifact, not evidence of learned phishing semantics. Benchmarking ten detectors across classic machine learning, Turkish/multilingual transformers, and an instruction-tuned large language model, eight exceed 0.99 macro F1 in distribution and are statistically indistinguishable (McNemar, p > 0.05), whereas out of topic TF-IDF linear models lead (logistic regression 0.947, 95% CI [0.926, 0.966]; 0.959 under the body-only protocol we recommend to users of the corpus) and mBERT collapses (0.570). Because persuasion phrasing is shared across themes, this protocol measures transfer to unseen theme vocabulary rather than robustness to novel phrasing. The transformer deficit out of topic is, for two of three encoders, a thresholding rather than a ranking failure: mBERT retains 0.960 out-of-topic AUC while its recall at the default 0.5 rule falls to 0.263. The corpus, splits, generation framework, and the full leakage audit, including its negative results, are available for research use.
PhishingGAT, a detector that fuses word-level semantic features with structural ones and is hardened against adversarial perturbation, is presented, a detector that fuses word-level semantic features with structural ones and is hardened against adversarial perturbation.
R. Kodali, Siva Rama Krishna T Dr· International Journal of Inn...· 0 citations
The outcome indicates that phishing detection improves when models incorporate attributes that match the target region, and the benefit is largest in emerging digital markets where attackers combine local social engineering with rapid infrastructure changes.
Jaber M. Al-Dulimi· Journal of Al-Turath Univers...· 0 citations
Background: WordPress plugins account for the large majority of disclosed CMS vulnerabilities, and learning-based detectors report high accuracy on synthetic corpora and random splits. Methods: We build a benchmark from 1757 real plugin CVEs (6666 indexed identifiers, 1281 plugins), yielding 30,860 labeled PHP function...
Z. Tashenova, Aisultan Aitmagambetuly, A. Urynbassarova et al.· Big Data and Cognitive Compu...· 0 citations
Source shift can make phishing detectors appear stronger while concealing unsafe class-specific errors. This study evaluated a three-class workflow (Legitimate, Suspicious, and Phishing) using source-aware governance, controlled augmentation, family-disjoint diagnostics, and pre-specified safety gates. The frozen train...
Akam Aziz, Umran Abdullah Haje· Indonesian Journal of Comput...· 0 citations
Phishing remains one of the most pervasive threats to Internet users, and email remains its predominant delivery channel. Email content is the attack surface of phishing: it is what the victim reads and what automated defenses inspect. Yet the composition of modern phishing content is poorly measured. Prior work has ch...
Jaehwan Park, Woonghee Lee, Fu-Jiao Ji et al.· 0 citations
This paper compares five classifiers: Multinomial Naive Bayes, Random Forest, Bidirectional Long Short-Term Memory, BiLSTM, DistilBERT, and BERT-base, and finds that BERT-base achieves the highest F1-score and DistilBERT the lowest, representing the strongest accuracy–latency trade-off in the evaluated environment.
Andre Sebastian Samaniego Buñay, Ariel Misael Orellana Albarracin, Joel Marcelo Chuquimarca Pomagualli· Enfoque UTE· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.