Skip to content
Open access

Comparative Evaluation of Prompt Engineering Configurations and Open-Source Large Language Models for Phishing Email Detection

2026 · International journal of research and innovation in social science · 0 citations

Abstract

Large Language Models (LLMs) offer a flexible approach to phishing-email classification, but their performance can vary with both prompt design and model selection. This study evaluates these factors through a two-stage experimental design using a balanced dataset of 1,000 emails comprising 500 phishing and 500 legitimate messages. In Stage 1, five complete prompt-engineering configurations —Zero-Shot, Few-Shot Demonstration, Expert Role-Based, Structured Criteria-Guided, and CoT-Style Structured Reasoning Prompting— were compared using Mistral. The best-performing configuration was then held constant in Stage 2 to compare five locally deployed open-source LLMs: Mistral, Phi-3, Qwen2.5, Llama 3.1, and Gemma 2. Performance was assessed using accuracy, precision, recall, F1-score, valid-output coverage, and processing time. Structured Criteria-Guided Prompting achieved the strongest prompt-level performance, with 86.33% accuracy and an F1-score of 85.43%. In the model comparison, Gemma 2 achieved the highest accuracy (87.00%), recall (90.00%), and F1-score (87.38%), while maintaining 100% valid-output coverage. Qwen2.5 achieved the highest precision (95.73%) and the shortest recorded processing time of 1 h 7 min 46 s. Valid-output coverage also varied across prompt configurations, with Expert Role-Based Prompting producing the lowest coverage at 88.2%. Overall, the results show that prompt configuration and model choice jointly influence phishing detection performance, false-alarm rates, output reliability, and local processing efficiency.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.