Hybrid Machine Learning Framework for Phishing Website Detection
Abstract
Phishing websites continue to be a major cybersecurity threat because attackers create deceptive web pages that imitate trusted banking, e-commerce, social media, and government platforms to collect sensitive user information. Traditional blacklist-based and rule-based detection methods are limited because they mainly identify known threats and require frequent manual updates. This paper proposes a hybrid machine learning framework for phishing website detection using character-level Term Frequency-Inverse Document Frequency (TF-IDF), Truncated Singular Value Decomposition (SVD), K-Means clustering, Random Forest, Extreme Gradient Boosting (XGBoost), and an Artificial Neural Network (ANN). In the proposed approach, TF-IDF and SVD generate compact URL representations, while K-Means cluster labels and centroid-distance features enhance the feature space before supervised classification. The prediction probabilities generated by Random Forest, XGBoost, and ANN are combined through a weighted soft-voting ensemble. Experimental results on a balanced URL dataset demonstrate that the hybrid model achieves strong classification performance with an accuracy of 93.76%, precision of 90.70%, recall of 80.55%, F1-score of 85.32%, and ROC-AUC of 97.65%. The proposed framework provides a practical and scalable approach for URL-based phishing detection and real-time web security applications.