Cross-Dataset Reliability of Random Forest Phishing URL Detection under Heterogeneous Feature Representations: Implications for Cybersecurity Decision-Making
Phishing URL detectors are commonly evaluated using training and test samples drawn from the same dataset, a practice that can overstate reliability when operational data differ in source, collection period, class composition or feature-generation process. This study evaluates Random Forest transferability under within...