Aug 2026· National Aerospace and Electronics Conference· pp. 460-465· 0 citations· 27 references
Abstract
Active Directory (AD) underpins enterprise identity management and is among the most heavily targeted assets in modern enterprise networks. When red teams emulate attackers targeting it, they typically chain BloodHound for relationship mapping, Impacket for protocol manipulation, and Responder for credential interception. Despite rapid advances in Large Language Models (LLMs), reliably automating this BIR workflow is still difficult, and the difficulty is not primarily one of reasoning. We analyze three sources of integration friction that block automation. The first is data dependency and contextual fragility, the second is temporal and procedural synchronization, and the third is semantic and operational inconsistency. We propose three architectural responses. They are tool wrappers, sidecar validators, and a shared Blackboard. Throughout, we carefully separate what we validate from what we propose. We empirically evaluate the sidecar pattern across four failure modes $(N=80)$. The wrapper and Blackboard layers are presented as architecture whose end-to-end validation remains future work. For the sidecar, deterministic pre-validation lets code, rather than the LLM, route around connectivity failures, with equivalent latency and only a modest token overhead in the structured output. The same experiment serves as a component-level ablation. It contrasts workflow behavior with and without the sidecar, converting ambiguous tool errors into structured, actionable directives.
AEGIS is presented, a policy enforcement component that enables administrators to define fine-grained safeguards against resource abuse across heterogeneous MCP tools and modalities and detects and mitigates abusive behaviors while preserving the flexibility of MCP-based agent ecosystems.
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidenc...
Feng-Peng Li, Qi-Zhou Wang, Yu-Ke Hu et al.· 0 citations
This work shows that a malicious developer can pair a benign-looking wrapper with crafted metadata to deterministically alter post-generation behavior without modifying model weights, training data, or inference backend, and introduces TIF-BAH, a lightweight middleware defense that verifies wrapper integrity and record...
Nokimul Hasan Arif, Qian Lou, Meng Zheng· 0 citations
CIPR (Coding In Poisoned Repos), the first benchmark that systematically varies PLCs in poisoned real-world repositories, is introduced and highlights that coding agent vulnerability is not a static property, but a dynamic outcome shaped by everyday user configurations.
Fu-Kang Zhu, Bin-Bin Zhao, Rui-Xiao Lin et al.· 0 citations
LLM agents that invoke external tools face critical safety vulnerabilities when malicious manipulations exploit their implicit trust in tool outputs and metadata. However, identifying these vulnerabilities through testing is challenging due to the need to bypass safety guardrails with semantically legitimate inputs, th...
Yu-Chen Shao, Zi-Qun Bao, Yu-Heng Huang et al.· 0 citations
It is argued that GraphQL constitutes a principled, testable alternative to function calling for agentic systems, combining lower cost, stronger safety, and improved cognitive robustness.
Viktor Zhakhalov· CEUR Workshop Proceedings, V...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.