A Framework for Real-Time Phishing Resource Locator Detection Using Structural and Lexical Feature Analysis Based on Machine Learning
Abstract
Phishing attacks remains a leading, rapidly evolving threat in cybersecurity domain, where cyber-criminals deploy fraudulent websites that are deceptive in nature acts exactly same as legitimate platforms to transfer sensitive user credentials, financial data and corporate data to an external location. Traditional counter measures including blacklist-based systems and rule-based filters demonstrates limited efficiency against newly created malicious URLs and real-time generated phishing URLs. This paper presents PhishGuard+, a random forest-based phishing detection framework that performs real-time suspicious URL classification through complete structural and lexical feature analysis of the URL. The proposed system utilizes a dataset of 10,000 labeled URLs, from which 48 distinct features are extracted that includes URL patterns, domain characteristics, and heuristic indicators. A Random Forest classifier is trained on preprocessed data comprises min-max normalization and the Synthetic Minority Oversampling Technique (SMOTE) to address skewed data distribution. The trained model is deployed with a Flask-based RESTful API, that real-time inference through an interactive web interface. Experimental evaluation demonstrates an overall classification accuracy of 98.4%, with precision and recall exceeding 98% across both phishing and legitimate URL categories. The proposed framework offers a computationally efficient, low latency lightweight detection solution deployable in browser extensions, email security gateways, and corporate cybersecurity infrastructures, providing robust early-warning capabilities without requiring full webpage content analysis.