Jul 2026· Annual International Computer Software and Applications Conference· pp. 2369-2374· 0 citations· 14 references
Abstract
The growing complexity and frequency of cyberattacks make cybersecurity risk assessment an increasingly demanding task for organisations, requiring substantial expertise, resources, and adherence to established standards. This work explores the applicability of Large Language Model (LLM) to cybersecurity risk assessment, with a focus on threat identification and risk scoring. The paper presents a standalone consistency analysis across five models, measuring accuracy and stability under lexical, structural, and noisy prompt perturbations using an OWASP-oriented rubric. Building on the analysis results, we present a modular LLM-based system that combines Retrieval-Augmented Generation, MITRE ATT&CK-Aligned threat evaluation, rubric-constrained risk scoring, and a Judge Reviewer, orchestrated through a Beliefs–Desires–Intentions control loop. The validation against incidents from the VERIS and EuRepoC datasets highlights limitations and weaknesses, and allows identifying the architectural and structural mitigations that can reduce prompt sensitivity in LLM-based risk assessment.
The findings show that LLMs can approximate structured cybersecurity reasoning under controlled representations, but do not apply it robustly, which has important implications for the design and evaluation of AI-assisted security decision-support systems.
To improve cybersecurity across industries, Cyber Threat Intelligence (CTI) is becoming increasingly crucial. This systematic review explores how CTI practices are evolving in response to advancements in Artificial Intelligence (AI), particularly in the context of Large Language Models (LLMs). We examined 61 peer-reviewed studies using the PRISMA methodology, which demonstrates a strict selection procedure founded on specified inclusion, exclusion, and quality standards. This approach aligns with the scope of similar systematic reviews in the field of cyber threat intelligence. The review provides a comparative synthesis of CTI research capabilities across threat detection and prediction, attribution, forecasting, and automated reporting. We classify these approaches into three categories: conventional methods, those enhanced by AI and Machine Learning, and those based on LLMs. Our findings indicate that LLMs offer significant advantages in contextual reasoning, processing unstructured threat intelligence, and generating actionable mitigation plans. However, challenges such as model explainability, data privacy, system interoperability, and standardization impede their integration into operational environments. In addition to highlighting the potential and practical limitations of LLMs in CTI, this study identifies research gaps and proposes methods to create scalable, secure, and flexible CTI systems that support real-time cyber defense.
Hilalah Alturkistani, Abdul Ghafar Jaafar, S. Chuprat et al.· International journal of res...· 0 citations
Large Language Models (LLMs) have rapidly evolved into powerful general-purpose systems with advanced natural language processing, code generation, and reasoning capabilities, leading to their increasing adoption in cybersecurity. However, their dual-use nature introduces both significant defensive opportunities and emerging offensive threats. This study presents a PRISMA-ScR-guided scoping review to systematically map the current landscape of LLM applications in cybersecurity, addressing their roles as both threat enablers and defensive tools while identifying key governance challenges and future research directions. Literature published between January 2017 and December 2024 was identified through structured searches of IEEE Xplore, ACM Digital Library, Scopus, Web of Science, and arXiv, supplemented by grey literature and citation snowballing. Studies were screened using predefined inclusion and exclusion criteria, and relevant information was extracted using a standardized data-charting framework followed by thematic narrative synthesis. The review synthesizes evidence from 153 eligible studies, demonstrating that LLMs substantially enhance offensive capabilities such as phishing, malware generation, vulnerability discovery, and adversarial attacks, while simultaneously improving defensive functions including threat detection, vulnerability management, incident response, security automation, and analyst support. The review further identifies critical limitations related to hallucinations, model reliability, privacy, misuse, and governance, highlighting the need for trustworthy deployment frameworks, standardized evaluation benchmarks, and robust regulatory safeguards. By integrating evidence across technical, operational, and governance perspectives, this scoping review provides a comprehensive evidence-based synthesis of the evolving role of LLMs in cybersecurity and outlines priorities for future research and responsible deployment.
: The rapid evolution of adversarial cyber threats demands proactive, scalable security testing methodologies capable of producing realistic, organization-specific attack scenarios. Conventional approaches, including manual red-teaming, scripted Breach and Attack Simulation (BAS) platforms, and tabletop exercises, are constrained by high expert dependency, limited scenario variability, and an inability to dynamically adapt to an organization’s unique threat profile. This paper proposes and evaluates a Large Language Model (LLM)-Assisted Threat-Driven Testing System that integrates the MITRE Adversarial Tactics, Techniques, and Common Knowledge (MITRE ATT&CK) framework v14, a structured knowledge base of adversarial tactics, techniques, and procedures (TTPs), with GPT-based language models accessed through the OpenAI API, to automate the generation of contextually tailored cyber-attack narratives. The system employs a service-oriented architecture implemented in Python, utilizing Streamlit for the interactive web interface, Pandas for ATT&CK data management, and LangChain as the prompt-orchestration middleware. Evaluation encompassed structured feedback surveys from 30 cybersecurity professionals representing security operations, red-teaming, and incident response roles, together with quantitative analysis using three performance metrics: ATT&CK Technique Coverage (ATC = 85%), False Positive Rate (FPR = 3.2%), and False Negative Rate (FNR = 11%). These results confirm that the system achieves high scenario fidelity, strong ATT&CK alignment, and a generation latency of 2–8 s per scenario. Practically, the framework enables security teams, particularly resource-constrained organizations lacking dedicated red-team capabilities, to conduct high-fidelity threat simulation exercises aligned with current adversarial TTPs, without specialized AI expertise, thereby strengthening organizational cyber-readiness at significantly lower cost than traditional security testing approaches.
Praise Emeka Nze, A. Ademuwagun, Muktar Bello et al.· Journal of Cyber Security· 0 citations
With the increasingly aggressive cyber threat landscape for governments, businesses, and institutions, as information and/or cybersecurity implementations are increasingly under scrutiny by regulators, it has been pointed out that governance failure is one of the major reasons for a weakened cybersecurity posture. A major component of Cyber/information security governance is the development, adoption, and implementation of a comprehensive information and/or cyber security policy document. The policy document must be in compliance with international or national standards and, if possible, with regulatory guidelines. However, it is often observed that policy documents are often incomplete with respect to industry standards or regulations and require revision when subjected to a thorough audit. Identifying the gaps between the controls and processes documented in the policy and those required in the regulations or standards necessitates extensive manual effort. The advent of Generative AI tools such as Large Language Models (LLMs) led to use of LLMs and Agentic AI tools to automate such compliance checks, as seen in a few research publications in recent times. However, such reported use of LLMs are experimented with high resource environments such as expensive GPUs and memory based servers. For smaller organizations such expensive compute platform may not be easily available. In this article, we benchmark the compliance checking tasks on LLMs that do not require GPU and high memory usage and the effectiveness of such resource constrained LLMs in compliance checking. Our experiments demonstrated that the low resource LLMs can provide good agreement/accuracy in compliance checking of policy documents against standards by experimenting with ISO 27002:2022 controls against multiple policy documents.
R. Negi, Rishika Jain, Soumyo V Chakarborty et al.· 0 citations
Academic research on securing Large Language Models (LLMs) in cybersecurity currently
exists in silos. To address this fragmentation, this study develops the 'Holistic Deployment Risk
Model' (HDRM) through a qualitative thematic synthesis of nine 'cornerstone' articles, selected
through purposive sampling to ensure a representative cross-section of technical, ethical, and
organizational perspectives, consolidating technical vulnerabilities, autonomous agentic risks,
and adversarial misuse with organizational governance and human elements to identify critical
'blind spots'. The model clusters associated risks into five holistic, interdependent layers -
Governance, Data and Privacy, Model Behavior, Operational Security, and Integrations and
Infrastructure - while illustrating how vulnerabilities can propagate across these interdependent
layers. To help demonstrate concrete applicability for regulatory readiness and real-world uses,
components of the model are qualitatively mapped to current AI risk management frameworks:
NIST AI RMF 1.0 and ISO/IEC 23894. While the model is currently theoretical and requires
further investigation to confirm its efficacy, it offers organizations an actionable checklist to
approach secure LLM deployment. It emphasizes the critical need for established governance
before technical implementation and the importance of human-in-the-loop management for highrisk workflows. Ultimately, the work highlights key areas of 'ethical security' necessary to
responsibly develop and manage AI systems, including the mitigation of probabilistic decisionmaking impacts such as bias, misinformation, and privacy violations.
Yosef Nethaniel Ozeri, Ilan Schreiber, Amir Schreiber· International journal of adv...· 0 citations