A AraCTI-NER is introduced, a dataset of 10,312 token-level annotated samples over eight STIX-inspired entity types built by an LLM-assisted pipeline seeded with authentic Arabic cybersecurity articles, structurally validated and rebalanced through targeted generation.
Abstract
Automated extraction of structured threat information from unstructured cyber threat intelligence (CTI) underpins modern security operations, yet the supporting machine learning resources are almost exclusively English: no annotated Arabic CTI named entity recognition (NER) corpus has been published. We introduce AraCTI-NER, a dataset of 10,312 token-level annotated samples (275,530 tokens; 42,360 entity spans) over eight STIX-inspired entity types, built by an LLM-assisted pipeline seeded with authentic Arabic cybersecurity articles, structurally validated and rebalanced through targeted generation. We benchmark seven encoders from three families (Arabic-specialized, English cybersecurity-adapted, and multilingual) over three seeds under strict entity-level metrics, and release a 408-sentence expert-audited test subset (ATS-gold) whose reliability is quantified by a second independent expert validation (inter-annotator agreement 0.878 entity-level F1). XLM-RoBERTa Large attains the best mean F1 (0.7603; 0.7674 on ATS-gold), with AraBERTv2 close behind (0.7491), while both English-only cybersecurity encoders fall to ≈0.63, a separation that holds across every seed and survives expert correction, with the ≈3-point F1 decrease from ATS-silver to ATS-gold concentrated in Vulnerability and TTP. On 350 doubly annotated sentences from authentic Arabic cyber-incident news, a shift in both provenance and register, the strongest model reaches F1 = 0.5429 against an inter-annotator F1 of 0.616. AraCTI-NER establishes the first reproducible baseline for Arabic CTI NER and identifies domain-adaptive Arabic cybersecurity pre-training as the highest-value next step.
Cyber Threat Intelligence (CTI) reports are the primary source of actionable threat knowledge; however, most high-value content is published in unstructured text, creating a significant bottleneck for operational threat response. Manual extraction of Indicators of Compromise (IoC), cybersecurity entities, and Tactics,...
An Bao Tran, Kiệt Nguyễn Văn, Duy The Phan· International Conference on...· 0 citations
CyberNER is introduced, a two-stage pipeline to solve the multi-type NER and alias canonicalization problem in APT CTI reports and achieves a Macro-F1 score of 0.853, outperforming all four baselines.
Unnamalai K, Suriakala M· International journal of com...· 0 citations
A manually constructed dataset of 150 English-language CTI reports, each represented as STIX 2.1 based graphs, provides a benchmark for CTI information extraction, knowledge-graph construction, incident analysis, and threat attribution and indicates that locally deployed LLMs can support human reviewers in identifying...
Dipshikha Das, Arnab Banik, Md. Shariful Islam et al.· arXiv.org· 0 citations
The threat posed by adversarial prompts to large language models is becoming harder to ignore. Problems including prompt injection, jailbreaking, phishing, and Unicode-based attacks are now widespread. Most existing solutions protect against only one threat type, operate in English only, and provide no explanation for...
Abdullah M. Abughallous, Somia Abufakher· IEEE Jordan Conference on Ap...· 0 citations
STINER, a taxonomy and expert-annotated corpus for extracting strategic intelligence from social media streams is introduced, and how social-media-driven extraction can surface early signals of the SafePay ransomware campaign prior to its retrospective characterization in vendor threat landscape reports is illustrated.
Yasir Ech-Chammakhy, Oussama Azrara, J. Chbili et al.· 0 citations
PAP_NER is presented, the first large-scale, gold-standard Vietnamese administrative NER corpus comprising 162,801 sentences with 205,807 entity annotations across five entity types critical for e-Government workflows, with implications for low-resource language NLP research.
Dinh-Dien La, Tien-Bang Tran, Ngoc-Huy Du et al.· PLoS ONE· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.