Skip to content
Open access

AraCTI-NER: A Dataset and Benchmark for Arabic Cyber Threat Intelligence Named Entity Recognition

Aug 2026 · Electronics · 0 citations · 40 references

TL;DR

A AraCTI-NER is introduced, a dataset of 10,312 token-level annotated samples over eight STIX-inspired entity types built by an LLM-assisted pipeline seeded with authentic Arabic cybersecurity articles, structurally validated and rebalanced through targeted generation.

Abstract

Automated extraction of structured threat information from unstructured cyber threat intelligence (CTI) underpins modern security operations, yet the supporting machine learning resources are almost exclusively English: no annotated Arabic CTI named entity recognition (NER) corpus has been published. We introduce AraCTI-NER, a dataset of 10,312 token-level annotated samples (275,530 tokens; 42,360 entity spans) over eight STIX-inspired entity types, built by an LLM-assisted pipeline seeded with authentic Arabic cybersecurity articles, structurally validated and rebalanced through targeted generation. We benchmark seven encoders from three families (Arabic-specialized, English cybersecurity-adapted, and multilingual) over three seeds under strict entity-level metrics, and release a 408-sentence expert-audited test subset (ATS-gold) whose reliability is quantified by a second independent expert validation (inter-annotator agreement 0.878 entity-level F1). XLM-RoBERTa Large attains the best mean F1 (0.7603; 0.7674 on ATS-gold), with AraBERTv2 close behind (0.7491), while both English-only cybersecurity encoders fall to ≈0.63, a separation that holds across every seed and survives expert correction, with the ≈3-point F1 decrease from ATS-silver to ATS-gold concentrated in Vulnerability and TTP. On 350 doubly annotated sentences from authentic Arabic cyber-incident news, a shift in both provenance and register, the strongest model reaches F1 = 0.5429 against an inter-annotator F1 of 0.616. AraCTI-NER establishes the first reproducible baseline for Arabic CTI NER and identifies domain-adaptive Arabic cybersecurity pre-training as the highest-value next step.

Read PDF

Similar papers

Conference Aug 2026

S-AECR: Word Intelligence and Explanation Strategy for Cyber Threat Intelligence Extraction and Mapping

Cyber Threat Intelligence (CTI) reports are the primary source of actionable threat knowledge; however, most high-value content is published in unstructured text, creating a significant bottleneck for operational threat response. Manual extraction of Indicators of Compromise (IoC), cybersecurity entities, and Tactics,...

An Bao Tran, Kiệt Nguyễn Văn, Duy The Phan · 0 citations
Review Jul 2026

A Structured Cyber Threat Intelligence Dataset Using STIX 2.1 Entities and MITRE ATT&CK Mappings

A manually constructed dataset of 150 English-language CTI reports, each represented as STIX 2.1 based graphs, provides a benchmark for CTI information extraction, knowledge-graph construction, incident analysis, and threat attribution and indicates that locally deployed LLMs can support human reviewers in identifying...

Dipshikha Das, Arnab Banik, Md. Shariful Islam et al. · 0 citations
Conference Jul 2026

SemGuard: A Triple-Anchor Semantic Security Gateway for Multilingual Prompt Attack Detection in Large Language Models

The threat posed by adversarial prompts to large language models is becoming harder to ignore. Problems including prompt injection, jailbreaking, phishing, and Unicode-based attacks are now widespread. Most existing solutions protect against only one threat type, operate in English only, and provide no explanation for...

Abdullah M. Abughallous, Somia Abufakher · 0 citations
Preprint Aug 2026

STINER: Automated Extraction of Strategic Cyber Threat Intelligence from X

STINER, a taxonomy and expert-annotated corpus for extracting strategic intelligence from social media streams is introduced, and how social-media-driven extraction can surface early signals of the SafePay ransomware campaign prior to its retrospective characterization in vendor threat landscape reports is illustrated.

Yasir Ech-Chammakhy, Oussama Azrara, J. Chbili et al. · 0 citations
Open access Jul 2026

PAP_NER: A large-scale vietnamese administrative named entity recognition corpus and hybrid deep learning architecture

PAP_NER is presented, the first large-scale, gold-standard Vietnamese administrative NER corpus comprising 162,801 sentences with 205,807 entity annotations across five entity types critical for e-Government workflows, with implications for low-resource language NLP research.

Dinh-Dien La, Tien-Bang Tran, Ngoc-Huy Du et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.