Skip to content

Leakage-Aware Cross-Dataset Evaluation of Prompt Injection Detection Using Classical Machine Learning and Transformer Models

Sep 2026 · Yalvaç akademi dergisi · Vol 11, pp. 102-131 · 0 citations · 25 references

TL;DR

A leak-aware cross-dataset evaluation framework is presented to examine the robustness of classical machine learning and Transformer-based prompt injection detection models in the face of data source changes, and shows that internal validation results are not enough for prompt injection detection.

Abstract

The widespread adoption of systems based on Large Language Models has made the reliable detection of prompt injection attacks a critical requirement. However, high performance achieved on training and test splits generated from the same data source does not guarantee that models can generalize to prompts from different sources. In this study, a leak-aware cross-dataset evaluation framework is presented to examine the robustness of classical machine learning and Transformer-based prompt injection detection models in the face of data source changes. During the data preparation process, empty records, duplicate prompts, conflicting labels, and text overlaps between datasets were checked. In this context, 38,184 duplicate records and 14 instances with conflicting labels were removed, and the 198 common prompts identified between the training and external test sets were removed only from the training set. Using WordHash and CharHash representations, SGD Logistic and Linear SVM models, as well as DistilBERT and DeBERTa-v3-small, were evaluated on an internal dataset consisting of 426,073 cleaned requests; the models were also tested on an independent dataset of 5,000 examples. While the models achieved performance in the range of approximately 0.997–1.000 in the internal evaluation, significant performance losses were observed in the external evaluation. DeBERTa-v3-small delivered the most balanced results, with an accuracy of 0.7360, a balanced accuracy of 0.7320, a Macro-F1 of 0.7284, and an attack sensitivity of 0.7119. The domain classifier achieved an ROC-AUC of 0.9770 hence indicating a remarkable shift in the distribution across data sources. The results show that internal validation results are not enough for prompt injection detection. Independent external validation, data leakage verification and domain shift analysis should be fundamental parts of reliable model evaluation.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Multi-Class, Multi-Tier Network Intrusion Detection: A Comprehensive and Reproducible Benchmark

Machine learning (ML) and deep learning (DL) have dominated Intrusion Detection System (IDS) research in recent years. Unfortunately, many existing studies have produced inflated results and unreliable benchmarks due to critical oversights and mistakes in the ML and DL pipeline, from data collection and labeling to fea...

Yu-Feng Xin, Bryant Goseland, Mohamed Rahouti · 0 citations
Open access 2026

Development of Hybrid Distilled Bidirectional Encoder Representations from Transformers and Support Vector Machine Model for Phishing Email Detection

Phishing email attacks are popular and common cyber security threats because they become more sophisticated and rely on social engineering techniques. The contextual and semantic features present in phishing emails are often missed by standalone machine-learning methods and transformer-based models like Bidirectional E...

D. Buhari, I. Ismaila, Noel et al. · 0 citations
Review Open access Aug 2026

Sequence-Aware Dataset Auditing for Leakage-Free Benchmarking of YOLO Detectors for Bottle Detection

This work presents a reproducible YOLO-based pipeline for bottle detection in sandy environments, emphasizing dataset integrity, leakage-free evaluation, and deployment-oriented model selection. A one-class dataset of 1585 images and 3167 annotated bottles was audited to identify annotation-format defects and near-dupl...

Rafael Reveles-Martínez, Sebastián Burciaga-Sosa, J. Celaya-Padilla et al. · 0 citations
Open access Oct 2026

Malware Detection by Machine Learning based on the LightGBM Model

Wannacry has grown to be considered one of the biggest threats affecting computer security over the past ten years, which has prompted continuous research on detection and mitigation techniques. The EMBER 2018 dataset, which contains sample WPE files categorized as harmless and harmful, is used to synthesize research f...

Amjad Jumaah Frhan · 0 citations
Open access 2026

MalBERT-Temporal: Transformer-Based Zero-Day Malware Detection in Windows Executable Binaries Under Strict Temporal Isolation

A Transformer-based malware detection approach named MalBERT-Temporal that reshapes the 2,381-dimensional BODMAS PE feature vector into 16 contiguous feature-group tokens and then processes these tokens with Transformer encoder layers employing multi-head self-attention to model interactions among all tokens is introdu...

Manar Alanazi, Israa Alsiyat · 0 citations
Conference Open access May 2026

A Calibrated and Explainable Bimodal Machine Learning Framework for Hybrid Intrusion Detection

Modern communication systems face critical gaps in detecting unknown attacks and rare threat classes due to extreme data imbalance and black-box decision logic. We propose a bimodal framework of calibrated and explainable machine learning (ML) for network security, unifying known-class precision with open-set generaliz...

Hafsa Aslam, Yue Li, Saba Aslam et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.