Skip to content

DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents

Sep 2026 · 0 citations · 51 references
Computer Science

TL;DR

Direct Relational Set-Risk Pruning is introduced, which formulates agent-history compression as risk-constrained selection over deletion sets and shows that decision-conditioned relations, retained-context information, pair interactions, and abstention each contribute to reliable pruning.

Abstract

Long-horizon language-model agents accumulate reasoning traces, tool exchanges, and observations whose relevance changes with the current decision. Existing compression strategies often score historical units independently, but the safety of deleting several units is generally not determined by their singleton scores: redundant evidence, accumulated small effects, and the information that remains after deletion all matter. We introduce Direct Relational Set-Risk Pruning (DRSR), which formulates agent-history compression as risk-constrained selection over deletion sets. Offline, DRSR constructs exact counterfactual supervision by jointly deleting protocol-valid history Blocks and measuring the change in teacher-forced likelihood of the same recorded next output. A lightweight scorer then predicts set-level harm from online-visible relations between candidate history and the current pre-action state, together with deleted-retained and pairwise set structure. At deployment, DRSR evaluates a small set of structurally valid deletion candidates with the lightweight scorer and removes the largest feasible set under recency, protocol, budget, and learned-risk constraints, abstaining when no set is sufficiently safe. On WorkBuddyBench Full260, DRSR increases mean reward from 0.699 to 0.802 while reducing total model tokens by 20.820%. On the fixed Eval40 comparison, it obtains 0.794 reward at 1.211M tokens per task, using 35.850% fewer tokens than the uncompressed agent. Mechanistic analyses and ablations further show that decision-conditioned relations, retained-context information, pair interactions, and abstention each contribute to reliable pruning.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes w...

Zi-Yang Yu, Liang Zhao, Bo-Wen Zhu et al. · 0 citations
Preprint Aug 2026

Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory

Experiments establish ReTree as an effective self-correcting memory abstraction for long-horizon search, and show that ReTree consistently outperforms Full-Trajectory ReAct in question-answering and search benchmarks.

Aijun Yang, Qianxue Guo, Ziyi Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajec...

Ming-Xuan Wang, Fei Luo, Bo Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

Retrieval assembles repository context by ranking passages for relevance to the current query. A coding agent halfway through an issue has already read much of what such a ranker returns. Relevance is scored per passage, but sufficiency belongs to the set: independently scored passages can fill the budget with support...

Zhe-Xi Feng, Rui-Yi Zhang, Yong-Bo Yang et al. · 0 citations
Preprint Aug 2026

Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents

Experiments on WebShop, ALFWorld, and visual Sokoban show consistent improvements across language and vision-language models, while diagnostic ablations support the effectiveness of Bellman fixed-point value estimation and show that step-level credit should be incorporated selectively rather than uniformly into the fin...

Hongxi Yan, Ziyue Huang, Shichao Fan et al. · 1 citation
#artificial intelligence Preprint Sep 2026

StateComp: Learning When to Compress History in Long Horizon Agents

Long-horizon agents continuously accumulate interaction history during task execution, yet the importance of past interactions changes as the agent state evolves. Existing context management methods largely compress history based on fixed windows, periodic schedules, or current relevance, overlooking a more fundamental...

Ming-Xuan Wang, Hong-Yue Chen, Ying-Long Guo et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.