Skip to content

Merging Cyber Threat Intelligence Through Retrieval-Augmented Generation and Small Language Models for Rich Threat Representation

Sep 2026 · 0 citations · 36 references
Computer Science

TL;DR

An automated pipeline is proposed that derives an actionable representation of a cyberattack from heterogeneous CTI sources and produces an enriched Attack Graph that captures a coarse, tactic-aligned progression of the attack and annotates each step with explicit pre-conditions and post-conditions, and an enriched description.

Abstract

Modern cybersecurity operations rely on CTI collected from heterogeneous sources, including semi-structured threat representations, IoCs, and narrative technical reports. However, these artifacts are often insufficient in isolation to reconstruct how an attack unfolds, under which conditions each step is feasible, and which traces it leaves behind. In practice, analysts must manually correlate partial evidence scattered across multiple and only partially structured sources, delaying the design of effective prevention, detection, and response actions. To address this gap, we propose an automated pipeline that derives an actionable representation of a cyberattack from heterogeneous CTI sources. The pipeline combines a RAG architecture with a locally deployable SLM, used to consolidate such evidence and infer missing operational details. Starting from a semi-structured threat representation and auxiliary CTI documents, the pipeline produces an enriched Attack Graph that captures a coarse, tactic-aligned progression of the attack and annotates each step with explicit pre-conditions and post-conditions, and an enriched description. This representation supports prevention by exposing execution requirements, detection by highlighting observable traces, and response by clarifying the temporal progression of the attack. Then, due to the lack of validated datasets with ground-truth information on the temporal evolution of real-world attacks, we test the complete pipeline on 10 real-world case studies spanning multiple threat types, including backdoors and staged downloaders delivered via phishing. A manual assessment across 10 real-world case studies provides initial evidence that the generated graphs are consistent with expected attack progressions, indicating that the proposed approach can support analysts by consolidating dispersed CTI evidence into a structured and actionable view of attacks.

View source

Similar papers

Conference Sep 2026

ATT&CK-Aware Summarization and Structuring of Cyber Threat Intelligence Reports using Large Language Models

Cyber Threat Intelligence (CTI) reports constitute a critical source of knowledge for Security Operations Centers (SOCs), providing information about adversarial behaviors, attack techniques, and emerging threats. However, these reports are typically unstructured, lengthy, and difficult to process efficiently in operat...

Santiago Cortés Ocaña, S. Rueda, Miguel Ángel Jimeno Paba · 0 citations
Preprint Aug 2026

STINER: Automated Extraction of Strategic Cyber Threat Intelligence from X

STINER, a taxonomy and expert-annotated corpus for extracting strategic intelligence from social media streams is introduced, and how social-media-driven extraction can surface early signals of the SafePay ransomware campaign prior to its retrospective characterization in vendor threat landscape reports is illustrated.

Yasir Ech-Chammakhy, Oussama Azrara, J. Chbili et al. · 0 citations
#natural language process... Preprint Aug 2026

BEACON: Behavior-Anchored Cross-Source Knowledge Graph Construction for Cyber Threat Intelligence

BEACON is an LLM-driven framework for cross-source CTI knowledge graph construction that constructs and releases two human-annotated datasets from 34 sources and outperforms all baselines by at least 23% and 9%, respectively.

Chang-Ze Li, Yutong Cheng, Tsania Camila Finnisa et al. · 0 citations
Preprint Aug 2026

The Anatomy of a Prompt Injection: A Component Model for Structured Analysis

This paper formalizes the structure of prompt-injection artifacts, enabling defenders, red teamers, and cyber threat intelligence (CTI) teams to label, compare, and mutate attacks without relying on fragile string matching.

Jeremy McHugh · 0 citations
Sep 2026

An adaptive framework for cyber threat recognition using contextualised embeddings and contrastive learning

A domain-adapted NER framework that combines a fine-tuned robustly optimised BERT pretraining approach encoder with a bidirectional long short-term memory layer to capture both long and short patterns in threat reports, demonstrating the framework’s value for real-world CTI analysis and automated cyber defense.

Bhubharv Mohan Sharma, Aruna Malik · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.