AI-Based Phishing Detection System Using Random Forest with Smote and Explainability
Abstract
One of the most widespread and destructive forms of Cyber Security breach is the Phishing Attack. Phishing attacks often defeat email filtering systems due to their heightened use of social engineering techniques that by-pass traditional filtering mechanisms. This research employs a Random Forest classifier, alongside TF-IDF based feature extraction methods, to classify Emails. The major innovation of this research is the implementation of SHAP (SHapley Additive exPlanations) technology in relation to the output of the Random Forest to provide explanations of each prediction in an understandable and user-friendly manner. Users may access this system via a simple copy-and-paste interface or uploading a file to submit email content for real-time analysis. The results of the analysis will consist of prediction (Phishing/Safe), confidence rating, level of risk, and a detailed listing of each risk factor as calculated using SHAP based explanations. The Random Forest model will have an accuracy of 96.12% with 94.4% precision and a recall rate of 95.87%, significantly better than the results generated using traditional rule-based methods. Response times for Email analysis averaged 3 to 5 sec. Combining high detection rates with explainable AI improved Cyber Security Awareness, while also offering practical and easily accessible protection from Phishing attacks for individual.