Sep 2026· Yalvaç akademi dergisi· Vol 11, pp. 102-131· 0 citations· 25 references
TL;DR
A leak-aware cross-dataset evaluation framework is presented to examine the robustness of classical machine learning and Transformer-based prompt injection detection models in the face of data source changes, and shows that internal validation results are not enough for prompt injection detection.
Abstract
The widespread adoption of systems based on Large Language Models has made the reliable detection of prompt injection attacks a critical requirement. However, high performance achieved on training and test splits generated from the same data source does not guarantee that models can generalize to prompts from different sources. In this study, a leak-aware cross-dataset evaluation framework is presented to examine the robustness of classical machine learning and Transformer-based prompt injection detection models in the face of data source changes. During the data preparation process, empty records, duplicate prompts, conflicting labels, and text overlaps between datasets were checked. In this context, 38,184 duplicate records and 14 instances with conflicting labels were removed, and the 198 common prompts identified between the training and external test sets were removed only from the training set. Using WordHash and CharHash representations, SGD Logistic and Linear SVM models, as well as DistilBERT and DeBERTa-v3-small, were evaluated on an internal dataset consisting of 426,073 cleaned requests; the models were also tested on an independent dataset of 5,000 examples. While the models achieved performance in the range of approximately 0.997–1.000 in the internal evaluation, significant performance losses were observed in the external evaluation. DeBERTa-v3-small delivered the most balanced results, with an accuracy of 0.7360, a balanced accuracy of 0.7320, a Macro-F1 of 0.7284, and an attack sensitivity of 0.7119. The domain classifier achieved an ROC-AUC of 0.9770 hence indicating a remarkable shift in the distribution across data sources. The results show that internal validation results are not enough for prompt injection detection. Independent external validation, data leakage verification and domain shift analysis should be fundamental parts of reliable model evaluation.
Machine learning (ML) and deep learning (DL) have dominated Intrusion Detection System (IDS) research in recent years. Unfortunately, many existing studies have produced inflated results and unreliable benchmarks due to critical oversights and mistakes in the ML and DL pipeline, from data collection and labeling to fea...
Phishing email attacks are popular and common cyber security threats because they become more sophisticated and rely on social engineering techniques. The contextual and semantic features present in phishing emails are often missed by standalone machine-learning methods and transformer-based models like Bidirectional E...
D. Buhari, I. Ismaila, Noel et al.· International journal of res...· 0 citations
This work presents a reproducible YOLO-based pipeline for bottle detection in sandy environments, emphasizing dataset integrity, leakage-free evaluation, and deployment-oriented model selection. A one-class dataset of 1585 images and 3167 annotated bottles was audited to identify annotation-format defects and near-dupl...
Rafael Reveles-Martínez, Sebastián Burciaga-Sosa, J. Celaya-Padilla et al.· Technologies· 0 citations
Wannacry has grown to be considered one of the biggest threats affecting computer security over the past ten years, which has prompted continuous research on detection and mitigation techniques. The EMBER 2018 dataset, which contains sample WPE files categorized as harmless and harmful, is used to synthesize research f...
Amjad Jumaah Frhan· WSEAS transactions on system...· 0 citations
A Transformer-based malware detection approach named MalBERT-Temporal that reshapes the 2,381-dimensional BODMAS PE feature vector into 16 contiguous feature-group tokens and then processes these tokens with Transformer encoder layers employing multi-head self-attention to model interactions among all tokens is introdu...
Manar Alanazi, Israa Alsiyat· International Journal of Adv...· 0 citations
Modern communication systems face critical gaps in detecting unknown attacks and rare threat classes due to extreme data imbalance and black-box decision logic. We propose a bimodal framework of calibrated and explainable machine learning (ML) for network security, unifying known-class precision with open-set generaliz...
Hafsa Aslam, Yue Li, Saba Aslam et al.· International Conference on...· 0 citations