Multi-Strategy RAG With ATT&CK STIX Metadata-Enriched Embeddings for Technique Extraction
Abstract
Extracting Tactics, Techniques, and Procedures (TTPs) from Cyber Threat Intelligence (CTI) reports and mapping them to the MITRE ATT&CK framework is a critical challenge in security operations. The abstract and compressed nature of ATT&CK technique descriptions makes automated mapping of attack techniques difficult. Existing zero-shot Large Language Model (LLM)-based approaches predict techniques without external knowledge, often resulting in low recall and high false positive rates. Existing Retrieval-Augmented Generation (RAG)-based approaches embed procedure text without context from source and technique metadata and rely on a single similarity search strategy, limiting embedding quality and retrieval diversity. In this paper, we propose a Multi-Strategy RAG framework with ATT&CK STIX Metadata-Enriched Embeddings, which enriches procedure embeddings with deterministic metadata fields from ATT&CK STIX Data and combines multiple retrieval strategies. We prefix the procedure descriptions of 14,998 ATT&CK STIX Data entries with source name, source type, and associated technique information prior to embedding, thereby incorporating STIX object and relationship metadata into the embedding space beyond simple textual similarity. Unlike existing contextual retrieval methods that rely on LLM-generated summaries, the proposed method utilizes metadata parsed directly from ATT&CK STIX Data, ensuring full reproducibility. Furthermore, we introduce a multi-strategy union retrieval that combines three retrieval strategies—similarity, diversity, and source type (malware, tool, and intrusion-set)—via set union, followed by a two-pass LLM verification mechanism that filters unsupported candidates while retaining a safety fallback. An ablation study across eight metadata combinations identifies source type as the single most critical factor determining retrieval performance. The proposed method achieves a micro F1 score of 92.46% on the CTI-ATE benchmark (46 malware descriptions, 292 ground-truth techniques), demonstrating consistent performance gains over zero-shot LLMs, single-strategy RAG, and multi-strategy ensemble baselines. Furthermore, we validate generalization on two additional public benchmarks: on the large-scale AthenaBench (500 scenarios, approximately $11 \times $ the scale of CTI-ATE) the framework attains 78.2% top-1 accuracy, surpassing the strongest zero-shot baseline (GPT-5, 76.0%), and on the long-form, noisy reports of AnnoCTR it reaches 60.47% micro F1, above the dataset’s reported inter-annotator agreement (approximately 54%). Across six Enterprise ATT&CK snapshots from v17.0 to v19.1, mean row retrieval coverage remains 98.02%; at generation temperature 1.0, the version-wise mean Micro F1 ranges are 94.02–94.30% for Gemma 4 31B IT QAT and 93.00–94.49% for GPT-OSS 120B. A 27-condition cross-dataset retrieval sweep and a stage-level error audit further quantify the coverage–context trade-off and the principal failure boundaries. In addition, a retrieval coverage of 98.0% is achieved at $K{=}15$ , suggesting that performance improvements can be achieved by incorporating STIX-derived metadata without additional model training.