TrustShiftProbe is introduced, an evaluation and defense framework with four contributions: a stateful temporal threat model of the agent-server lifecycle as a benign conditioning phase followed by an adversarial defection at a trust horizon, and a language-agnostic attack engine that instantiates each variant as a compromised MCP server across four production domains.
Abstract
The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model agents to external tool backends. This openness introduces a severe server-side threat we term TrustShift: a compromised MCP server behaves benignly during an initial conditioning phase, building operational reliance and suppressing agent skepticism, before switching to an adversarial payload once an interaction threshold is reached. The evasion is temporal, not syntactic: benign at deploy time, the server's defection is invisible to predeployment static analysis, which sees only the honest phase. Switched payloads range from overt structural violations to schema-valid manipulations, the latter preserving outer protocol compliance to evade runtime middleware filters. Crucially, TrustShift originates in the server-controlled tool channel, not user prompts (unlike indirect prompt injection) or the transport (unlike man-in-the-middle): the adversary is the trusted server endpoint itself. We introduce TrustShiftProbe, an evaluation and defense framework with four contributions: (1) a stateful temporal threat model of the agent-server lifecycle as a benign conditioning phase followed by an adversarial defection at a trust horizon; (2) a language-agnostic attack engine that instantiates each variant as a compromised MCP server across four production domains; (3) SHIELD, a multi-tier, zero-oracle runtime defense at the MCP transport boundary that audits server payloads against behavioral baselines learned during clean trust windows; and (4) a taxonomy of nine TrustShift variants spanning three execution mechanisms (structural violation, semantic corruption, scope expansion) and three adversarial objectives (disruption, exfiltration, and their combination). Across frontier proprietary and open-weight models, TrustShift attacks achieve a 69.5% mean attack success rate that SHIELD mitigates to 42.7%.
The Model Context Protocol (MCP) enables LLMs to invoke external tools, but every tool interaction exposes the model to attacker-controlled text through multiple input channels (tool descriptions, tool results, sampling messages) that share a single context window without privilege separation. In this paper, we present...
Confidential MCP is presented, a set of backward-compatible extensions to MCP that enable standardized, auditable tool calling within and across TEE boundaries and introduces a three-zone enclave-partitioned server topology, a programmable Anonymization Transform Layer (ATL) with formal parameter classification and ent...
Ankur Aggarwal· International journal of com...· 0 citations
The Explainable Adaptive Zero Trust Framework (EAZTF) is introduced, a cloud-native security layer that continuously reevaluates the legitimacy of API actions throughout a session and is evaluated against four adversarial evasion strategies.
O. Singh, Yagyaraj Pandey, Nandini Pathak· 0 citations
Agentic Provenance (AgentProv), the first action-based identity audit for agentic LLM APIs, is introduced: AgentProv fingerprints a deployed model through its categorical tool-call distribution and decides identity via an MMD permutation test.
Xun Wang, Bihe Zhao, Michael Backes et al.· 1 citation
Experimental results show that although current agents can reliably uncover the problems exposed by alerts, they struggle to proactively investigate the disk for silent intrusions and to produce comprehensive, verified remediation plans, with no model achieving complete detection and remediation on any single range.
Le-Han Wang, Boli Chen, Ruixue Ding et al.· arXiv.org· 0 citations
The increasing number of vulnerabilities in operating systems, together with sophisticated kernel-level threats (e.g., rootkits), has weakened the effectiveness of traditional in-kernel protection mechanisms. Since these defenses operate at the same privilege level as the kernel, they share the same attack surface and...
Zhen-Ling Duan, Pan Dong, Renshuang Jiang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.