Skip to content
Preprint

OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents

Aug 2026 · 0 citations · 20 references
Computer Science

TL;DR

OBLIVION, a controlled benchmark and defense harness for revoked-skill resurrection, and results support workflow-level evaluation beyond checking explicit skill entries support workflow-level evaluation beyond checking explicit skill entries.

Abstract

Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift creates a new skill revocation problem: after a skill is removed from an explicit registry, an agent may still reconstruct it from residual carriers such as archives, transcripts, schemas, or memory entries. We study this problem as operational skill unlearning, where the goal is not parameter-level forgetting, but preventing a deployed agent from rebuilding a revoked skill through primitive tools. We introduce OBLIVION, a controlled benchmark and defense harness for revoked-skill resurrection. OBLIVION models each episode as a source-to-sink workflow, applies Cross-Surface Coherent Erasure to reduce residual carriers, and uses frozen workflow remediation near dangerous sinks. On the locked 88 attack episodes, the no-defense arm reaches formal attack success rate 1.0. OBLIVION reduces the rate to 0.114 and impact-weighted exposure to 0.115 while keeping locked utility at 1.0 and benign block rate at 0. In a separate skill-attack-derived sandbox, OBLIVION reduces attack success from 1.0 to 0.2 and impact-weighted exposure from 1.0 to 0.213 while preserving all utility controls. These results support workflow-level evaluation beyond checking explicit skill entries.

View source

Similar papers

Preprint Aug 2026

Repo2Skill-Evo: Repository Skills Go Stale in Silence

Repo2Skill-Evo casts each release transition as a skill-maintenance task: given a V1 skill set and the official V1-to-V2 patch, an agent must update obsolete skill content while preserving guidance that remains valid.

Chenyuan Duan, Ge Shi, Zineng Mao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents

Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memory, pending tool plans, and, under every serving API, a KV cache. Yet today's"forget"operations delete a plaintext memory record and stop, leaving every artifact derived from the revoked information intact. We f...

Chao Yao, Yangbo Wei, Zhen Huang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents

Persistent memory is entering production-oriented agent platforms to help long-horizon agents accumulate experience across sessions. Yet a post-task curator agent restricted to completed trajectories can preserve errors, overgeneralize partial evidence, or retain stale knowledge. We introduce environment-probing curati...

Susheel Suresh, Hazel Mak, Sahil Bhatnagar et al. · 0 citations
Preprint Aug 2026

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

StarHarness offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks by stratifying tasks according to baseline failure behavior, separating proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization.

Esakkivel Esakkiraja, D. Akhiyarov, Vikas Yadav et al. · 1 citation
Preprint Aug 2026

SkillJack: Persistent Skill Backdoors in Self-Evolving Agents

Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning, such attacks only affect agents when poisoned records are retrieved as context. We uncover a new and more fundamental risk: poisoned experie...

Zong-Hao Ying, Xiang-Fan Wu, Hui-Yu Wu et al. · 3 citations
Preprint Aug 2026

ContextWeave: A Real-World Workflow Benchmark

ContextWeave is introduced, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams and motivates memory systems that optimize not only retrieval relevance but also reliable use during execution.

Bo Wang, Yu Yao, Enxi Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.