Cross-Dataset Reliability of Random Forest Phishing URL Detection under Heterogeneous Feature Representations: Implications for Cybersecurity Decision-Making
Abstract
Phishing URL detectors are commonly evaluated using training and test samples drawn from the same dataset, a practice that can overstate reliability when operational data differ in source, collection period, class composition or feature-generation process. This study evaluates Random Forest transferability under within-dataset and heterogeneous cross-dataset conditions and considers the implications of transfer failure for cybersecurity decision-making. A 6,000-sample source dataset was constructed from PhishTank phishing URLs and an Alexa-derived legitimate URL list, while 6,000 records from the UCI Phishing Websites dataset formed an unseen structured target set. A semantic compatibility audit distinguished eight high-equivalence URL-reproducible features from three limited-equivalence measurements and 19 non-transferable attributes. Three Random Forest configurations were evaluated using grouped 10-fold cross-validation, source hold-out testing, registered-domain-grouped validation and one-way target transfer. RF-200 achieved 0.752 accuracy and 0.799 balanced accuracy on the source hold-out set; registered-domain grouping reduced balanced accuracy to 0.779. On the target dataset, the 11-feature model collapsed to an all-phishing prediction pattern, producing 0.633 accuracy and 1.000 phishing recall but 0.000 specificity and 0.500 balanced accuracy. Removing SSLfinal_State and Abnormal_URL produced a 9-feature sensitivity representation with 0.590 balanced accuracy, although one limited-equivalence port variable remained. The results show that aggregate accuracy can conceal complete class-level failure and that feature-generation semantics materially affect transferability. For operational cybersecurity, the findings emphasize independent target evaluation, domain-aware validation, transparent feature definitions, class-specific metrics and cautious use of model outputs in automated blocking or warning decisions.