A Framework for Phishing Email Detection with Statistical Text Representation and Semantic Feature Analysis
Abstract
Phishing emails are still a major cybersecurity risk because of more sophisticated social engineering methods. Many traditional detection techniques are losing effectiveness. The increasing complexity and variety of spoofing attacks require better observation systems. Within this research, a comparative foundation is proposed to compare TF-IDF and domain-specific NLP feature extraction methods for detecting phishing emails with machine learning models. Multiple classifiers, including Logistic Regression, Random Forest, Extra Trees, DSNLP-(LR, RF, ET), and TF-(LR, RF, ET), are trained and tested on a curated dataset that has optimized feature selection. According to experimental findings, the TF-Extra Trees model achieves a peak accuracy of 99.19%. It outperforms other baseline models and current methods found in recent studies. The use of statistical and specialized text features performs efficiently. The proposed approach provides an excellent, scalable, and reliable way to improve phishing email testing systems.