Skip to content

PLACEMEM: Toward a Compute-Aware Memory Plane for Lifelong Agents

Jul 2026 · arXiv.org · Vol abs/2607.04089 · 0 citations · 7 references
Computer Science

TL;DR

This work presents PLACEMEM as a systems position on lifelong-agent memory, instantiated by an executable control-plane prototype that demonstrates correction-aware control-plane behavior today and a concrete roadmap for replay-aware serving integration in future lifelong-agent systems.

Abstract

Lifelong agents need more than larger context windows and better retrieval. They need memories that can persist, evolve, and be corrected without forcing the serving stack to recompute the same history on every turn or silently reuse stale runtime state. We present PLACEMEM as a systems position on lifelong-agent memory, instantiated by an executable control-plane prototype. The central claim is that agent memory should be represented as versioned capsules that unify semantics, provenance, validity, and reusable runtime state under one correction-aware identity. In the current prototype, capsules drive prompt-level text retrieval, KV-aware routing, and cascading invalidation over live streamed backends; prospective layer-frontier replay is intentionally framed as a deeper integration agenda rather than a claimed engine feature. We describe a vLLM-first prototype with persistent capsule state, concurrency-safe invalidation, an OpenAI-compatible routing sidecar, a typed metadata contract, and a benchmark harness that measures live first-token latency, reuse, and post-correction behavior. The result is both an executable artifact that demonstrates correction-aware control-plane behavior today and a concrete roadmap for replay-aware serving integration in future lifelong-agent systems.

View source

Similar papers

Preprint Aug 2026

Context as an Environment: Programmatic Context Management for Long-Horizon Agents

LLM agents increasingly take on long-running tasks whose history grows far beyond a single model context window. Existing approaches compress earlier interactions or extract selected information into fixed memory representations, committing to what to preserve before future needs are known. We present Scroll, a context...

Yin Lin, Elaine Ang, E. Zhu et al. · 2 citations
Jul 2026

ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control

A write-time admission gate that, before committing a candidate fact m extracted from context c, queries the LLM K times for a soft support score and admits m only when the average exceeds a threshold, and reduces to a single forward pass in a log-probability variant for latency-sensitive deployments.

Yan Zhang, Shibo Li · 2 citations
#machine learning Preprint Sep 2026

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

Modern LLM agents operate in persistent workspaces whose accumulated history can exceed both GPU KV capacity and the model's native context window. Existing systems typically compact older context into summaries or retrieve it later as text, either losing fine-grained execution evidence or repeatedly prefilling content...

Di Chai, Leye Wang, Ze-Shen Su et al. · 0 citations
Jul 2026

MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory

MemTxn is a governance layer outside the answer model that verifies whether an update is supported by its source and restores the application-visible state after a fault, and achieves the highest average F1 across all twelve answer-model configurations.

Han-Shuai Cui, Zhiqing Tang, Z. Yao et al. · 2 citations
Jul 2026

ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory

ChronoMem is the first open-source system and benchmark for systematic semantic global memory rollback in LLM agents, and a post-exposure evaluation protocol that tests whether an agent can behave counterfactually after rollback by answering queries and summarizing history as if future updates had never occurred.

Yongye Su, Wujiang Xu, Chaoji Zuo et al. · 1 citation
Preprint Sep 2026

UNISON: A Co-Designed Near-Memory Scheduler of Session KV Residency for LLM Agents

Large language models are increasingly composed into agent loops that plan, call tools, and resume the same task after each action. These loops press a shared memory hierarchy harder than conventional multi-turn chat, because they hold a growing key-value (KV) prefix across tool waits and place many sessions on one SRA...

Fan He, Yan Li, Xiao-Yang Zeng · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.