Aug 2026· Pragmatic Cybersecurity· 0 citations· 38 references
Computer Science
TL;DR
The solution, LSABRE, is a multi-LLM framework that improves robustness across various attacks, maintaining 86% detection accuracy even under strong adaptive adversarial attacks, and proposes a robust multi-LLM defense architecture designed to preserve detection reliability under adaptive adversarial conditions.
Abstract
The rise of social media bots poses a persistent threat, enabling misinformation, public opinion manipulation, and erosion of trust in online platforms. To combat this, machine learning systems have been developed to detect and limit bot activity. However, attackers continuously adapt through adversarial optimization, behavior imitation, and semantic manipulation strategies, creating an escalating arms race with detection tools. Recent advances in LLMs have significantly improved bot detection by enabling deeper semantic and contextual analysis. However, this shift also introduces new attack surfaces, allowing adversaries to craft exploits that directly target LLM reasoning and generation mechanisms. Industry tools like Anthropic’s Claude Code Security similarly leverage LLMs for security, motivating our study of their attack surfaces. In this work, we explore both offensive and defensive aspects of LLM-powered, threat-specific cybersecurity applications. While centered on the challenge of social media bot detection, our methodology and insights generalize to a broad class of LLM-powered cybersecurity systems, including phishing detection, email classification, fraud analysis, and more. We introduce two novel adversarial attack strategies that systematically exploit semantic and contextual weaknesses of LLM-based classifiers, degrading LLM performance in bot detection by up to 48%, and propose a robust multi-LLM defense architecture designed to preserve detection reliability under adaptive adversarial conditions. Our solution, LSABRE, is a multi-LLM framework that improves robustness across various attacks, maintaining 86% detection accuracy even under strong adaptive adversarial attacks.
In recent years, phishing attacks have grown exponentially in scale, frequency, and sophistication, placing a significant burden on organizations and security personnel. Attackers leverage advanced obfuscation techniques to evade detection systems and ensure their emails reach users’ inboxes. Additionally, attackers di...
Email spam and phishing attacks remain a critical security threat. Adversaries increasingly exploit large language models to craft contextually convincing malicious messages, and existing spam detection systems often struggle to keep pace. Generalization across diverse and evolving attack scenarios is limited, which re...
Experimental results show that LLMs, when guided by rubric-based prompts and supplemented with ATT&CK domain knowledge, achieve robust performance across detection, localization, and TTP mapping tasks.
Joon-Young Gwak, Aubrey Strier, Zhaohan Xi et al.· 1 citation
This paper systematically synthesizes 27 studies of attacks against MGTD and the available evidence on corresponding defenses, categorizing existing research into four major types of evasion strategies: watermark attacks, paraphrasing attacks, prompt-based attacks, and adversarial-text attacks, and summarizes the avail...
De-Yu Meng, Tad Gonsalves· Neural Networks· 0 citations
This work provides the first study of such a whole-system defense, especially with respect to a deployed and operational capability, and shows an increase in product abuse coverage, a 30% reduction in monthly alerts, and adaptability to changes in malicious actors'behavior.
Shaefer Drew, Michael Brautbar, Paul Knight et al.· 0 citations
A high-efficiency detection framework utilizing DistilBERT, a distilled knowledge representation of the BERT transformer is proposed, substantiate the viability of Knowledge Distillation as a mechanism to deploy state-of-the-art semantic security filters on edge infrastructure.
Mrinal Mrinal, Neeraj Kumar· International Journal of Cre...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.