Skip to content
Preprint

Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

TTP-R1 is proposed, a two-stage framework that combines retrieval-augmented supervised fine-tuning (SFT) with reinforcement learning using verifiable rewards (RLVR) and applies Group Relative Policy Optimization with a decomposed reward that directly supervises the precision, recall, and output format of the predicted technique set.

Abstract

Mapping cyber threat intelligence (CTI) text to MITRE ATT&CK techniques is essential for structured threat analysis, yet manual annotation is costly and does not scale. The ATT&CK taxonomy comprises several hundred attack techniques, and a single CTI passage may describe multiple techniques, making accurate and complete extraction challenging. Existing automated approaches fall short in different ways: multi-label classifiers struggle with severe class imbalance and the large label space, while LLM-based methods--retrieval pipelines and fine-tuned generators--optimize token-level objectives that treat technique annotation as sequence generation rather than set prediction, lacking direct supervision on whether the predicted technique set is correct and complete. We propose TTP-R1, a two-stage framework that combines retrieval-augmented supervised fine-tuning (SFT) with reinforcement learning using verifiable rewards (RLVR). A hybrid retriever first narrows the large label space to a candidate set, and a fine-tuned LLM learns to select the correct techniques. We then apply Group Relative Policy Optimization with a decomposed reward that directly supervises the precision, recall, and output format of the predicted technique set. Across four CTI benchmarks, TTP-R1 achieves the best average F1, improving sub-technique-level F1 by 7.4 percentage points over Claude Sonnet 4.5 with retrieval augmentation, while running 28x faster when served as an 8B-parameter model on a single GPU.

View source

Similar papers

Sep 2026

An adaptive framework for cyber threat recognition using contextualised embeddings and contrastive learning

A domain-adapted NER framework that combines a fine-tuned robustly optimised BERT pretraining approach encoder with a bidirectional long short-term memory layer to capture both long and short patterns in threat reports, demonstrating the framework’s value for real-world CTI analysis and automated cyber defense.

Bhubharv Mohan Sharma, Aruna Malik · 0 citations
Open access 2026

Multi-Strategy RAG With ATT&CK STIX Metadata-Enriched Embeddings for Technique Extraction

Extracting Tactics, Techniques, and Procedures (TTPs) from Cyber Threat Intelligence (CTI) reports and mapping them to the MITRE ATT&CK framework is a critical challenge in security operations. The abstract and compressed nature of ATT&CK technique descriptions makes automated mapping of attack techniques difficult. Ex...

Ji-Min Lee, K. Kim, Kwonmoo You et al. · 0 citations
Open access Aug 2026

Automated MITRE ATT&CK Technique Classification Using OSINT and Advanced NLP

This paper offers an automated solution to the problem of categorizing the threat descriptions based on OSINT into the MITRE ATT&CK techniques with a sophisticated model based on transformer and natural language processing, leading to more effective CTI automation and decision support to the security operations of a pr...

Ramesh Kumar Sharma, Dharmendra Kumar Singh, Rajesh P. Barnwal · 0 citations
#small language model Preprint Sep 2026

Merging Cyber Threat Intelligence Through Retrieval-Augmented Generation and Small Language Models for Rich Threat Representation

An automated pipeline is proposed that derives an actionable representation of a cyberattack from heterogeneous CTI sources and produces an enriched Attack Graph that captures a coarse, tactic-aligned progression of the attack and annotates each step with explicit pre-conditions and post-conditions, and an enriched des...

Nicola Deidda, L. Regano, Alessandro Sanna et al. · 0 citations
Preprint Aug 2026

LLMs for Zero-Shot Threat Detection via Structured Risk Indicators

It is shown that the quality of the generated risk indicators is the main driver of zero-shot cyber threat detection performance, and that retrieval mainly benefits weaker LLMs by generating more discriminative risk indicators, whereas stronger models achieve comparable performance without retrieved context.

A. Al-Ghamdi, S. Layeghy, Marius Portmann · 0 citations
Conference Sep 2026

ATT&CK-Aware Summarization and Structuring of Cyber Threat Intelligence Reports using Large Language Models

Cyber Threat Intelligence (CTI) reports constitute a critical source of knowledge for Security Operations Centers (SOCs), providing information about adversarial behaviors, attack techniques, and emerging threats. However, these reports are typically unstructured, lengthy, and difficult to process efficiently in operat...

Santiago Cortés Ocaña, S. Rueda, Miguel Ángel Jimeno Paba · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.