Comparative Evaluation of Prompt Engineering Configurations and Open-Source Large Language Models for Phishing Email Detection
Abstract
Large Language Models (LLMs) offer a flexible approach to phishing-email classification, but their performance can vary with both prompt design and model selection. This study evaluates these factors through a two-stage experimental design using a balanced dataset of 1,000 emails comprising 500 phishing and 500 legitimate messages. In Stage 1, five complete prompt-engineering configurations —Zero-Shot, Few-Shot Demonstration, Expert Role-Based, Structured Criteria-Guided, and CoT-Style Structured Reasoning Prompting— were compared using Mistral. The best-performing configuration was then held constant in Stage 2 to compare five locally deployed open-source LLMs: Mistral, Phi-3, Qwen2.5, Llama 3.1, and Gemma 2. Performance was assessed using accuracy, precision, recall, F1-score, valid-output coverage, and processing time. Structured Criteria-Guided Prompting achieved the strongest prompt-level performance, with 86.33% accuracy and an F1-score of 85.43%. In the model comparison, Gemma 2 achieved the highest accuracy (87.00%), recall (90.00%), and F1-score (87.38%), while maintaining 100% valid-output coverage. Qwen2.5 achieved the highest precision (95.73%) and the shortest recorded processing time of 1 h 7 min 46 s. Valid-output coverage also varied across prompt configurations, with Expert Role-Based Prompting producing the lowest coverage at 88.2%. Overall, the results show that prompt configuration and model choice jointly influence phishing detection performance, false-alarm rates, output reliability, and local processing efficiency.