Attacker-Controlled Inference: How Context Poisoning, Speculative Disclosure, Memory Persistence, and Privilege Hierarchy Failures Converge on a Candidate Framework for LLM Agent Runtime Exploitation
Abstract
Version 3 (2026-09-26). Version 3 corrects a misattributed citation and several overstated readings, found by an independent audit of version 2 and by a second review of the corrected draft. (1) The figures for context-based adversarial attacks on code generators (10.7x, from 3.5% to 37.4%; 100% on GPT-3.5-Turbo; 60-100% cross-model transfer; 2,800 experiments) and the phrase "systemic architectural vulnerabilities rather than model-specific flaws" were credited to arXiv:2606.10860. They come from arXiv:2606.10945 (Del Orbe, Hastings and Vaidyan), and they are now cited to it. arXiv:2606.10860 is still cited, only for the instruction-hierarchy (GW-DPO) material it contains. (2) An 81% detection rate is no longer read as "19% of attacks succeed". (3) A deployment claim is no longer attributed to the conceptual AOS paper (arXiv:2606.01508). (4) The claim that tool responses are "architecturally indistinguishable" from system messages is replaced with the source's own statement (arXiv:2606.02240). (5) Membrane (arXiv:2606.05743) is described as the jailbreak guardrail it is, not as a filter for an agent's external memory. (6) The delegation-observability result (arXiv:2606.09692) is stated within its own scope. (7) The shared-mechanism bridge is narrowed to the three injection classes. Speculative tool-call disclosure is an outbound leak, not an injection, and partial counter-evidence from arXiv:2606.10860 is noted. (8) An unsourced ranking of the classes by severity has been removed. The thesis is otherwise unchanged. The file has a neutral name, and the full list of corrections is at the top of the PDF. Version 2 was revised in response to an external structural review and an automated critique pass; its change log is kept as the "Response to Review" appendix in the PDF. Large language model (LLM) agents are increasingly deployed with live credentials, persistent memory, external tool access, and real-world actuation capability operating without continuous human oversight. This paper argues, as a heuristic reading, not a derivation, that four independently studied attack classes converge on a single structural property: the LLM agent's runtime context is the primary attack surface. For three of the classes, every layer of the stack that can write to that context is a potential injection point; for the fourth, speculative disclosure, the context is the source of an outbound leak rather than an injection target. The four classes are: (1) adversarial context poisoning through fabricated evidence and instruction hierarchy violations arXiv:2606.06244 arXiv:2606.10945 arXiv:2606.10860; (2) speculative tool-call disclosure, where intent leaks before any commit decision arXiv:2606.02483; (3) multimodal memory poisoning that achieves persistent, goal-agnostic re-execution arXiv:2606.10742; and (4) agent-skill supply-chain injection, where third-party skill packages serve as both code and instruction arXiv:2606.07131 arXiv:2606.01494. A fifth finding, that evaluation methodology artifacts (padding convention, split protocol) can produce a 67-fold false-alarm-rate swing in intrusion detection benchmarks, is discussed in a weakly-connected addendum as a candidate warning, not an established parallel, for LLM agent defense evaluation arXiv:2606.11098. Taken together, these findings suggest that securing LLM agents may require treating the runtime context window as a trust boundary with first-class enforcement semantics, not as an incidental input buffer. This remains a structural hypothesis, not an empirical result. The primary falsification path is explicit: a controlled experiment deploying all four attack classes simultaneously against a single agent stack with layered defenses, measuring whether any defense composition closes all four channels without introducing unacceptable over-refusal. Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted synthesis from an arXiv preprint corpus, originally drafted 2026-06-11, produced under the direction of Cristian Ruvalcaba, the accountable human author. Not peer-reviewed. Cited arXiv preprints: arXiv:2606.01494, arXiv:2606.01508, arXiv:2606.01567, arXiv:2606.02240, arXiv:2606.02483, arXiv:2606.05743, arXiv:2606.06244, arXiv:2606.07131, arXiv:2606.09692, arXiv:2606.10742, arXiv:2606.10749, arXiv:2606.10860, arXiv:2606.10945 (added in v3), arXiv:2606.11098 AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.