TrustTrace: Provenance-Aware Tracing of Indirect Prompt Injection Propagation in LLM-Assisted Learning Workflows
Abstract
Indirect prompt injection can persist through retrieval, generation, and revision after the visible payload disappears. Existing defenses primarily detect suspicious inputs or constrain immediate actions, but they provide less evidence about which source influenced which final-artifact span through a temporally valid route. This paper presents TrustTrace, a provenance-aware framework for source-to-artifact tracing in LLM-assisted learning workflows. TrustTrace constructs a heterogeneous temporal graph over sources, retrieved contexts, prompts, responses, revision operations, and artifact spans. A dual-flow encoder separates legitimate evidence transmission from unauthorized control transmission, while differentiable reachability supports source attribution, contaminated-span localization, activation-stage estimation, and complete-path recovery. A bounded counterfactual verifier performs fact-preserving neutralization, suffix replay, calibrated set-valued attribution, and abstention. The primary pipeline uses procedurally separated GPT-5.5 profiles; construct validity is therefore tested with independent human ratings, a Qwen3-235B-A22B-Instruct generation set, and an independent Qwen3 neutralization-and-judging stack. On the independently audited EduTrace-IPI subset, TrustTrace improves strict Complete-Path F1@3 from 0.631 to 0.684 over the architecture-matched terminal-supervision baseline and from 0.659 to 0.716 after disagreement-flagged workflows are excluded. Human ratings of 800 stratified artifact spans correlate with the continuous contamination score at Spearman $\rho =0.78$ , and an independent Qwen3 verifier preserves most of the primary result (Span-F1 0.731 and Complete-Path F1@3 0.671). On Qwen3-generated workflows, TrustTrace obtains Span-F1 0.681 and Complete-Path F1@3 0.598 without retraining. Binary detection gains remain comparatively small, and specialized guards remain stronger on selected native input- or action-level metrics. The evidence supports TrustTrace as a complementary forensic layer for calibrated span and path tracing rather than a replacement for preventive defenses.