Skip to content
Open access

Cross-Dataset Reliability of Random Forest Phishing URL Detection under Heterogeneous Feature Representations: Implications for Cybersecurity Decision-Making

2026 · International journal of research and innovation in social science · 0 citations

Abstract

Phishing URL detectors are commonly evaluated using training and test samples drawn from the same dataset, a practice that can overstate reliability when operational data differ in source, collection period, class composition or feature-generation process. This study evaluates Random Forest transferability under within-dataset and heterogeneous cross-dataset conditions and considers the implications of transfer failure for cybersecurity decision-making. A 6,000-sample source dataset was constructed from PhishTank phishing URLs and an Alexa-derived legitimate URL list, while 6,000 records from the UCI Phishing Websites dataset formed an unseen structured target set. A semantic compatibility audit distinguished eight high-equivalence URL-reproducible features from three limited-equivalence measurements and 19 non-transferable attributes. Three Random Forest configurations were evaluated using grouped 10-fold cross-validation, source hold-out testing, registered-domain-grouped validation and one-way target transfer. RF-200 achieved 0.752 accuracy and 0.799 balanced accuracy on the source hold-out set; registered-domain grouping reduced balanced accuracy to 0.779. On the target dataset, the 11-feature model collapsed to an all-phishing prediction pattern, producing 0.633 accuracy and 1.000 phishing recall but 0.000 specificity and 0.500 balanced accuracy. Removing SSLfinal_State and Abnormal_URL produced a 9-feature sensitivity representation with 0.590 balanced accuracy, although one limited-equivalence port variable remained. The results show that aggregate accuracy can conceal complete class-level failure and that feature-generation semantics materially affect transferability. For operational cybersecurity, the findings emphasize independent target evaluation, domain-aware validation, transparent feature definitions, class-specific metrics and cautious use of model outputs in automated blocking or warning decisions.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.