Skip to content
Preprint

LLMs for Zero-Shot Threat Detection via Structured Risk Indicators

Aug 2026 · 0 citations · 22 references
Computer Science

TL;DR

It is shown that the quality of the generated risk indicators is the main driver of zero-shot cyber threat detection performance, and that retrieval mainly benefits weaker LLMs by generating more discriminative risk indicators, whereas stronger models achieve comparable performance without retrieved context.

Abstract

We propose a two-stage large language model (LLM) framework for zero-shot detection of insider threats and advanced persistent threats (APTs) from heterogeneous security logs. The framework models user activity as chronological timelines and incorporates retrieval-augmented generation (RAG) to provide personalised behavioural context from each user's historical activity. Rather than performing end-to-end classification directly from raw logs, it first generates structured, interpretable sets of threat-specific risk indicators, which are then classified jointly across temporal sequences to capture attack patterns spanning multiple windows.The framework is evaluated on two benchmark datasets, CERT r5.2 for insider threat detection and PicoDomain for APT detection, using four combinations of two open-weight LLMs under both retrieval and non-retrieval settings. All configurations outperform the previous state-of-the-art LLM-based framework (GABM), with the best configuration improving the F1-score by 11.40 percentage points on CERT r5.2 and 31.50 percentage points on PicoDomain. Results further show that retrieval mainly benefits weaker LLMs by generating more discriminative risk indicators, whereas stronger models achieve comparable performance without retrieved context. The most effective assignment of LLMs to the two stages depends on the dataset. These findings show that the quality of the generated risk indicators is the main driver of zero-shot cyber threat detection performance.

View source

Similar papers

LADE: LLM-Assisted Advanced Persistent Threat Detection and Explanation

Experimental results show that LLMs, when guided by rubric-based prompts and supplemented with ATT&CK domain knowledge, achieve robust performance across detection, localization, and TTP mapping tasks.

Joon-Young Gwak, Aubrey Strier, Zhaohan Xi et al. · 0 citations
Review Aug 2026

ThreatLens: Evidence-Guided Ranking of High-Priority CVEs

Security teams must prioritize vulnerabilities before exploitation evidence is complete. Existing signals, such as CVSS, EPSS, advisories, and public exploits, are useful but fragmented and time-sensitive; retrospective rankings can therefore overstate performance by using evidence unavailable at decision time. We pres...

Soroush Motamedi Sedeh, Panteha Shahrivar, M. Qureshi et al. · 0 citations
Open access 2026

MLBRS: A Multi-Layer Behavioural Risk Scoring Framework for Insider Threat Detection

MLRS, a multi-layer behavioural risk scoring framework that combines rule-based scoring, statistical deviation analysis, and Isolation Forest-based anomaly detection to generate continuous employee-level risk scores, is presented.

V. L. Kartheek, A. Shah, Rishav Jain et al. · 0 citations
Open access Aug 2026

Tri-Level Network Attack Risk Stratification Using IDS-Contextual Ensemble Learning

An Intrusion Detection System (IDS)-contextual ensemble learning framework that assigns network traffic to three operationally meaningful risk tiers: High, Medium and Low is presented.

Reeta Mishra, Neelu Chaudhary · 0 citations
Review Jul 2026

Traceable LLM Reasoning for Fake-Order Fraud Detection

Results show that DeepScrub improves fraud review accuracy, reduces first-stage review workload, and provides traceable evidence for production risk-review workflows, showing that domain adaptation can matter more than model scale in this setting.

Si-Qi You, Bingsong Xu, Zhixiang Zheng et al. · 0 citations
Open access Aug 2026

Real-Time Detection and Mitigation of Prompt Injection Attacks in LLM-Integrated Enterprise Systems

Large language models (LLMs) embedded in enterprise workflows cannot structurally distinguish legitimate instructions from adversarial ones in the same token stream, making prompt injection OWASP's top LLM risk for two consecutive editions a persistent threat across direct and indirect vectors. This paper presents Prom...

Fatimah Alhamzawi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.