Machine-learning and DNN-based Email Phishing Recognition Methods
Abstract
Email spam and phishing remain one of the threats in the information era that harms information and financial security. This study, based on machine learning methods, uses a Kaggle dataset containing over 250,000 samples to classify emails into three categories: legitimate, phishing, and spam. A pre-processing pipeline was used, utilizing TF-IDF vectorization. Three models are evaluated: Complement Naive Bayes (CNB), Random Forest (RF), and Deep Neural Network (DNN). Results show that CNB achieved an overall accuracy of 89.6%, while RF demonstrated perfect precision (1.00) at the cost of low recall (0.55). The DNN model outperformed the two baselines, achieving an accuracy of 97.4%. Word analysis later showed characteristic patterns, such as the timestamps and dates in phishing emails and product-oriented vocabulary in spam. These results show that a complex machine learning model with careful training would excel at jobs like email blocking and provide a foundation for email security systems. Future work should explore Transformer-based architectures and hybrid feature engineering approaches to enhance robustness.