Skip to content
Open access

Predicting Emerging Cyber Threats Using Natural Language Processing (NLP)

Sep 2026 · Al-Noor Journal of Engineering Management and Computer Science · 0 citations · 9 references

Abstract

The rapid evolution of cyber-threats, including Advanced Persistent Threats (APTs) and dynamic malware strains, has made signature-based defense techniques obsolete. Proactive Cyber Threat Intelligence (CTI) and autonomous host-based detection have emerged as critical paradigms to circumvent modern threats before any system is exploited. This paper proposes a full-fledged, multi-level framework that approaches security telemetry and unstructured textual logs like natural language constructs. We evaluate our system on a heterogeneous dataset (10,000 samples containing 8 different threat domains: adware, botnet, phishing, ransomware, rootkit, spyware, Trojan, and worm) and introduce a dual-engine architecture: an Early-Stage Structural Identifier with static metadata feature engineering, using together an ensemble bagging tree classifier as predictor, and a Linguistic Context Predictor supported by Bidirectional Long Short-Term Memory (Bi-LSTM) networks and customized Transformers. In an empirical study, this was revealed to perform best when the structural metadata, which has high cross-family structural entropy, produced an accuracy baseline of 23.15%, while the linguistic engine that processes text-mined security intents reached an accuracy rate of up to 100.00% with a false positive rate of 0.00%. This leads to a scalable approach toward modern enterprise ecosystems that support resilient, real-time autonomous threat mitigation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.