Skip to content
Open access

A Comparative Evaluation of Large Language Models for Named Entity Recognition in Cyber Threat Intelligence

2026 · International Journal of Advanced Computer Science and Applications · 0 citations · 14 references

Abstract

Cyber Threat Intelligence reports combine analytical prose with dense technical indicators, making structured entity extraction a challenging but operationally valuable task. This study presents a comparative evaluation of three large language models – Claude Sonnet 4.6, GPT-5.4, and LLaMA 4 Scout – on a manu-ally annotated corpus of 21 real-world CTI reports across 15 entity types and 1284 ground truth instances. This study evalu-ates zero-shot and few-shot prompting conditions and studies the effect of iterative prompt refinement, focusing on explicit format constraints for cryptographic hash entities. Results show that Claude Sonnet 4.6 and GPT-5.4 achieve comparable perfor-mance under zero-shot conditions, with LLaMA 4 Scout trailing by a substantial margin. Few-shot prompting consistently reduc-es hallucination rates, but yields mixed F1 results, with exemplar cardinality emerging as a critical and underappreciated design factor. Entity extraction difficulty varies substantially across types, with technical indicator categories showing near-perfect performance and semantic categories such as tool and target sector posing the greatest challenges across all evaluated mod-els.

Read PDF