TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers
TrustShiftProbe is introduced, an evaluation and defense framework with four contributions: a stateful temporal threat model of the agent-server lifecycle as a benign conditioning phase followed by an adversarial defection at a trust horizon, and a language-agnostic attack engine that instantiates each variant as a com...