ATT&CK-Aware Summarization and Structuring of Cyber Threat Intelligence Reports using Large Language Models
Abstract
Cyber Threat Intelligence (CTI) reports constitute a critical source of knowledge for Security Operations Centers (SOCs), providing information about adversarial behaviors, attack techniques, and emerging threats. However, these reports are typically unstructured, lengthy, and difficult to process efficiently in operational environments. This paper presents an ATT&CKaware framework based on Large Language Models (LLMs) for the automatic extraction, structuring, and summarization of CTI reports. The proposed approach combines LLM-based knowledge extraction, semantic retrieval, and contextualization through the MITRE ATT&CK framework to transform unstructured threat intelligence into operationally useful representations. The framework extracts relevant information related to threat actors, ATT&CK tactics and techniques, affected assets, and potential impacts, generating structured outputs and SOC-oriented summaries. The proposed approach was evaluated using 553 CTI evidence instances associated with ten threat actors documented in MITRE ATT&CK. From this dataset, 200 samples were selected for experimental evaluation. Results show an average Exact Technique Accuracy of 43% and a Parent Technique Accuracy of 61%, indicating that LLMs are generally capable of identifying the correct ATT&CK behavioral category, although finegrained attribution of specific ATT&CK sub-techniques remains challenging. The findings demonstrate the feasibility of using LLMs to support CTI analysis and ATT&CK-oriented knowledge structuring, facilitating the integration of threat intelligence into SOC workflows.