Skip to content

Speculate with Memory: Lossless Acceleration for LLM Agents

Jul 2026 · arXiv.org · Vol abs/2607.12236 · 2 citations · ⚡ 1 influential · 36 references
Computer Science

TL;DR

Memory-augmented speculation yields a 19--39\% relative accuracy improvement on action prediction and up to a $2.5\times$ increase on observation prediction tasks with repetitive action spaces as memory accumulates and generalize across speculator models of varying cost.

Abstract

Speculative execution accelerates LLM agents by using a smaller, cheaper model to predict and pre-launch the next step while the environment is idle. However, existing speculators are stateless and discard all information between tasks, preventing prediction quality from improving with experience. We equip the speculator with three online memory systems that learn from past agent trajectories: a contrastive transition table tracking action-sequence statistics, an episodic memory retrieving contextually similar segments, and a confusion tracker suppressing recurring errors. We evaluate this approach on six benchmarks spanning three speculation types: action prediction, observation prediction, and chained prediction. Memory-augmented speculation yields a 19--39\% relative accuracy improvement on action prediction and up to a $2.5\times$ increase on observation prediction tasks with repetitive action spaces. These gains grow continuously as memory accumulates and generalize across speculator models of varying cost. All speculation is lossless because it runs during idle time at zero added wall-clock cost, and the actor's trajectory is identical to non-speculative execution.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

FlowState: Execution State as Memory for Long-Horizon LLM Agents

Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $\tau^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption.

Ming-Hao Li, Bang-Yan Li, Zi-Fan Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

StateComp: Learning When to Compress History in Long Horizon Agents

This work proposes State Conditioned Compression (StateComp), a framework that determines when historical interactions can be safely compressed according to the current agent state, and demonstrates that it reduces total agent and summarization tokens while maintaining task performance, and achieves a 12.67-fold speedu...

Ming-Xuan Wang, Hong-Yue Chen, Ying-Long Guo et al. · 0 citations
Preprint Aug 2026

Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents

The results suggest that effective long-horizon agent memory depends less on storing more information than on deciding which information should remain active, and that effective long-horizon agent memory depends less on storing more information than on deciding which information should remain active.

Q. Dao, Purvi Kathalkar, Kenneth Eaton · 1 citation
Preprint Sep 2026

AReaL-TIK: Stateful Agentic Optimization of Unified RL Kernels through an Optimization IR

Reinforcement learning (RL) post-training often uses distinct GPU kernels for rollout and policy update. In synchronous PPO and GRPO, numerical disagreement can perturb ratios between current token probabilities and those assigned during rollout. Recomputing rollout log-probabilities with the policy-update backend avoi...

Ran Yan, You-He Jiang, Jia-Yi Nie et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound

LAM is proposed, a Lossy Agent Memory system with three components: a deterministic deduplication rule with a substitution bound on retrieval scores - a bound on score perturbation, not a certificate of unchanged ranking; a memory manager that preserves the cached prefix and overlaps compaction with inference; and a pe...

Bai-Xi Sun, Le Chen, Anjir Ahmed Chowdhury et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.