Skip to content

Supply-Chain Trust at Runtime: How Agent Skill Injection, Inference-System Fingerprinting, Distributed Attack Evasion, and Speculative Disclosure Define a New Execution-Time Threat Topology

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Version 3 (2026-09-26) corrects a citation error in versions 1 and 2 and a false correction added in version 2, found by an independent audit and confirmed by re-checking every cited arXiv abstract. The inference-system fingerprinting findings, one of the four topics named in the title, were cited to arXiv:2605.29960; they come from arXiv:2605.29979 (Wimbauer et al., "Fingerprinting Inference Systems of Large Language Models"), which is now cited. arXiv:2605.29960 is the MemPoison memory-poisoning paper (Wang et al.); version 2 wrongly withdrew that correct attribution and published an editorial note and a citation integrity note saying it could not be verified. Both notes are deleted and the MemPoison attribution is restored. Version 3 also corrects the description of SkillSpector (a registry-side semantic scanner, not a runtime signal) and withdraws the argument that the ClawHub data show build-time and runtime signals to be orthogonal; describes the 75.8% MalSkillBench figure as a benchmark verification yield rather than an attack statistic; withdraws the unsupported claims that runtime integrity is the dominant unresolved problem and that the attacks have high success rates; corrects overstated wording about WebMCP and A2A; repairs three malformed citations of arXiv:2605.31593; and corrects the source count. The full list of corrections is at the top of the PDF, and the version 2 change log remains as an appendix with its erroneous citation paragraph marked as withdrawn. A follow-up review on the same day led to further fixes: distributed attacks are described as largely evading per-transcript monitors rather than being invisible to them, the shared structure is renamed observation before commitment (the source's term), the claim that post-hoc filtering is insufficient is limited to the speculative tool-call paper, an unsupported remark about well-aligned models is removed, and the Section 3.4 falsification test now uses a no-speculation baseline. Conventional software supply-chain security focuses on build-time artifacts: package hashes, signed binaries, dependency manifests. The emergence of LLM-based agentic systems introduces a parallel but distinct threat surface that operates at runtime, after the artifact has been delivered and before the user sees an output. This paper synthesizes five specific findings from recent cs.CR preprints to argue that runtime execution integrity is a distinct problem that build-time provenance controls do not address, and that it deserves to be treated as its own layer of agentic AI security. The thesis, stated plainly: an attacker who cannot compromise a model's weights or a package's source code can still target the execution context (the inference system, the tool registry, the memory pipeline, the communication graph, and the speculative dispatch channel), and several recent preprints document such attacks evading the monitors, filters or detectors they were tested against. This is a heuristic reading that unifies mechanistically distinct attacks under a shared structural property (execution-context manipulation), not a derivation from a single formal model, and it does not claim that these attacks are uniformly successful or that this is the most important problem in agentic AI security. The corpus sources span agent security (cs.CR), with findings drawn from: a runtime-verified benchmark of malicious agent skills in which the strongest skill-specific detector reaches 98.4% recall on code injection but collapses on prompt-injection and agent-control attacks [arXiv:2606.07131]; an inference-system fingerprinting study showing that hardware and software stack identity leaks through prompt-response deviations [arXiv:2605.29979]; a stateful monitoring study demonstrating that distributed agent attacks evade per-transcript monitors while completing hard cybersecurity tasks [arXiv:2605.31593]; a speculative tool-call privacy study showing that ghost tool calls irreversibly disclose user intent before commit [arXiv:2606.02483]; and a mid-session tool injection study identifying race-condition and lifecycle attacks in the WebMCP protocol [arXiv:2606.06387]. The primary falsification path: if a single-layer execution monitor that combines inference-deviation signatures, tool-registration provenance, and cross-account clustering achieves <1% false positive rate at >90% attack detection across all five attack classes simultaneously, the claim that these threats require separate runtime controls is falsified. Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted synthesis from an arXiv preprint corpus; version 1 drafted 2026-06-09, version 3 revised 2026-09-26. Cited arXiv preprints: 2605.29960, 2605.29979, 2605.31593, 2606.01494, 2606.01508, 2606.02240, 2606.02483, 2606.06387, 2606.07131, 2606.07150 AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18

Related blog posts

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.