Skip to content
Open access

Automated MITRE ATT&CK Technique Classification Using OSINT and Advanced NLP

Aug 2026 · International journal of computer information systems and industrial management applications · 0 citations

TL;DR

This paper offers an automated solution to the problem of categorizing the threat descriptions based on OSINT into the MITRE ATT&CK techniques with a sophisticated model based on transformer and natural language processing, leading to more effective CTI automation and decision support to the security operations of a practitioner.

Abstract

Open-Source Intelligence (OSINT) can be considered a crucial part of the present-day Cyber Threat Intelligence (CTI) due to delivering prompt information about the adversary activity using publicly accessible reporting and analysis. Nonetheless, the conversion of unstructured OSINT stories into structured forms like the MITRE ATT&CK model is a highly manual and subjective task. The proposed paper offers an automated solution to the problem of categorizing the threat descriptions based on OSINT into the MITRE ATT&CK techniques with a sophisticated model based on transformer and natural language processing. The suggested framework has combined OSINT preprocessing, threat behavior extraction, semantic representation learning and multi-label ATT&CK techniques classification with confidence-aware outputs. Large-scale experiments on a wide OSINT corpus show that the proposed method is far more effective compared to the ones that rely on keyword parameters and conventional machine learning baselines, especially when there are missing and imprecise threat specifications. Findings indicate greater accuracy, retrieval, and strength of a vast variety of ATT&CK methods, including categories of low density. This work is a step in the right direction by facilitating scalable and standardized ATT&CK mapping of noisy OSINT data, thus leading to more effective CTI automation and decision support to the security operations of a practitioner.

Read PDF

Similar papers

Jul 2026

Integrating FRACAS and FMECA with Natural Language Processing (NLP): An AI-Assisted Approach to Reliability Analysis

An algorithm is developed that automates and streamlines the analysis of equipment field-failure reports and other unstructured maintenance records and reduces the resource-intensive manual work required to prepare, interpret, and process FRACAS reports, thus enabling timelier, data-driven equipment reliability analysi...

Esther Yu, Guangjiang Cao, Y. Khalil et al. · 0 citations
Open access Jul 2026

Comparative Performance of AI-Generated Fake News Detection Pipelines on Romanian News Content

Fake news detection has become a major research topic at the intersection of artificial intelligence, data mining, and information security. In this paper, we evaluate the performance of English-trained algorithms on English translations of Romanian-sourced news articles, using a translation-mediated cross-domain evalu...

C. Coman, Costel Marian Dalban, Vlad Bătrânu-Pințea et al. · 0 citations
Open access Sep 2026

Comparative Analysis of Transformer-Based and ClassicalMachine Learning Models for Phishing Email Detection:A Multi-Source Dataset Evaluation with Explainability

This paper compares five classifiers: Multinomial Naive Bayes, Random Forest, Bidirectional Long Short-Term Memory, BiLSTM, DistilBERT, and BERT-base, and finds that BERT-base achieves the highest F1-score and DistilBERT the lowest, representing the strongest accuracy–latency trade-off in the evaluated environment.

Andre Sebastian Samaniego Buñay, Ariel Misael Orellana Albarracin, Joel Marcelo Chuquimarca Pomagualli · 0 citations
Conference Jul 2026

SemGuard: A Triple-Anchor Semantic Security Gateway for Multilingual Prompt Attack Detection in Large Language Models

The threat posed by adversarial prompts to large language models is becoming harder to ignore. Problems including prompt injection, jailbreaking, phishing, and Unicode-based attacks are now widespread. Most existing solutions protect against only one threat type, operate in English only, and provide no explanation for...

Abdullah M. Abughallous, Somia Abufakher · 0 citations
Preprint Aug 2026

Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence

TTP-R1 is proposed, a two-stage framework that combines retrieval-augmented supervised fine-tuning (SFT) with reinforcement learning using verifiable rewards (RLVR) and applies Group Relative Policy Optimization with a decomposed reward that directly supervises the precision, recall, and output format of the predicted...

Jiayun Zhang, Junshen Xu, Zejun Xie et al. · 0 citations
Review Open access Sep 2026

Hybrid LaBSE Semantic and Handcrafted Feature Fusion with Machine Learning for Fake Review Detection in Roman Marathi Code-Mixed Text

The findings demonstrate the effectiveness of combining multilingual semantic information with explicit linguistic and contextual characteristics for identifying deceptive reviews in Roman Marathi code-mixed environments and establish an initial benchmark for fake review detection in this low-resource setting.

Swapnil S. Nehar, R. Keole, Pravin P. Karde · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.