2026· Journal of Advances in Information Technology· Vol 17, pp. 1310-1320· 0 citations· 23 references
TL;DR
The results indicate that cross-dataset evaluation is essential for the generalization performance of a cross-dataset measure for cybercrime detection across network traffic, phishing email, and malicious Uniform Resource Locator domains.
Abstract
—Proposed cybercrime detection models demonstrate satisfactory performance on specific benchmark datasets; however, they are not always robust across diverse, practical scenarios. This paper examines the generalization performance of a cross-dataset measure for cybercrime detection across network traffic, phishing email, and malicious Uniform Resource Locator (URL) domains. The work combines three popular security benchmarks Canadian Institute for Cybersecurity Intrusion Detection System Dataset2017 (CIC-IDS2017), Canadian Institute for Cybersecurity Phishing Dataset2019 (CIC-Phishing2019), and Malicious URL 2020 in a single preprocessing and learning pipeline to avoid dataset-specific bias. One or more datasets are used to train models, which are then directly evaluated on previously unseen datasets to verify transferability across distribution shifts. We will discuss performance while considering accuracy, F1 − Score, Receiver Operating Characteristic (ROC), and fold-to-fold stability. Experiments demonstrate that direct training with a single random source results in significant performance deterioration, with a 23% decrease in F1 − Score when applied to unseen datasets. In contrast, the degradation caused by the proposed framework is kept below 10% and maintains Receiver Operating Characteristic–Area Under the Curve (ROC–AUC) values consistently above 0.90. Paired significance testing demonstrates that the gains in robustness are highly significant ( p < 0.01). The results indicate that cross-dataset evaluation is essential for the
This work contributes a comparative evaluation of lightweight detection models, consistent attention to security-critical metrics, and interpretable insights to support practical cybersecurity deployment.
Nwagbara Chisom Telvin, Gilbert Imuetin Osaze Aimufua, Raymond Ternenge Igbudu· FUDMA Journal of Sciences· 0 citations
Phishing URL detectors are commonly evaluated using training and test samples drawn from the same dataset, a practice that can overstate reliability when operational data differ in source, collection period, class composition or feature-generation process. This study evaluates Random Forest transferability under within...
N. Bahaman, Sarah Aqilah Arun, Erman Hamid et al.· International journal of res...· 0 citations
Phishing attacks remains a leading, rapidly evolving threat in cybersecurity domain, where cyber-criminals deploy fraudulent websites that are deceptive in nature acts exactly same as legitimate platforms to transfer sensitive user credentials, financial data and corporate data to an external location. Traditional coun...
P. Srivastava, Akhilesh Singh, Amit Virmani et al.· International Journal of Sci...· 0 citations
The increasing prevalence of encrypted network traffic has reduced the effectiveness of traditional intrusion detection methods that rely on payload inspection and signature-based analysis. Machine learning–based anomaly detection offers a promising alternative, but its effectiveness depends heavily on the representati...
Mark Marrero, Yan-Zhen Qu· European Journal of Electric...· 0 citations
GAFPNet (Generalization-Aware and False-Positive Controlled Framework Network), a five-module stacked-ensemble framework, is introduced as a generalization-aware and deployment-oriented phishing URL classifier for real-time filtering applications.
Mohammed Elias Basha S., M. N.· Journal of Trends in Compute...· 0 citations
: Phishing remains one of the most persistent attack vectors in cybersecurity, and the gradient-boosting models that now dominate its automated detection are frequently deployed as opaque classifiers, which limits analyst trust and slows incident response. This paper presents a cross-dataset, explainability-driven eval...
S. M. Nihal Ahmed, Afrim Hossen Khan, Mahmudur Rashid· Journal of Cyber Security· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.