Skip to content
Open access

An Explainable Ensemble Artificial Intelligence Framework for Real-Time Phishing Website Detection

Aug 2026 · International Journal of Computer Science and Mathematical Theory · 0 citations

Abstract

Phishing attacks remain a major cybersecurity threat, with the Anti-Phishing Working Group (APWG) recording 989,123 attacks in the fourth quarter of 2024 alone. Existing anti-phishing solutions are constrained by high false positive rates, reliance on static blacklists that cannot detect new phishing sites, and a lack of explainability in classification decisions. This study developed and evaluated an Explainable Ensemble Artificial Intelligence Framework for Real Time Phishing Website Detection. The proposed framework combined three base classifiers operating in parallel — a character-level 1D Convolutional Neural Network (CNN), a Random Forest (RF), and an Extreme Gradient Boost (XGBoost) — whose outputs were combined by a Logistic Regression meta-learner using a five-fold Out-of-Fold cross-validation strategy. Twenty URL features were extracted across Lexical, Character-based, Domain-based, and Binary categories. The framework was trained and evaluated on a stratified sample of 50,000 URLs from the Kaggle Phishing Site URLs dataset, split 70/15/15 for training, validation, and testing. SHapley Additive exPlanations (SHAP) were integrated to provide feature-level justification for every classification decision. The proposed ensemble achieved 97.48% accuracy, 98.43% precision, 96.51% recall, 97.46% F1-Score, and AUC-ROC of 0.9956, outperforming all individual base classifiers. The false positive rate of 2.43% directly addresses the primary weakness of existing systems. Suspicious Keywords Score, Special Character Ratio, and Domain Entropy Score were identified as the three most discriminating features. The framework was deployed as a real-time desktop application, correctly classifying a phishing URL at 99.7% confidence.

Read PDF