Skip to content
Preprint

LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference

Aug 2026 · 1 citation · ⚡ 1 influential · 65 references
Computer Science

TL;DR

An intrinsic memory method, LiveMem, is introduced, which augments a pretrained full-attention LLM with a memory state that preserves the historical information over the whole lifecycle while the main attention path retains a bounded KV window and establishes state continuity as a distinct and complementary abstraction for continual LLM inference.

Abstract

Long-running assistants and agents consume interaction streams that eventually outgrow the context. Existing context retention, summarization, and retrieval preserve access to selected history, but do not provide a persistent state over the full lifecycle when working context changes. We formulate this missing inference capability as \emph{state continuity under context turnover}: carrying computation forward through a fixed-capacity memory state whose lifetime is independent of the active context. We introduce an intrinsic memory method, \textbf{LiveMem}, which augments a pretrained full-attention LLM with a memory state that preserves the historical information over the whole lifecycle while the main attention path retains a bounded KV window. Context turnover and memory state maintaining, memory-oriented post-training, and state-aware serving jointly make this memory state load bearing after its originating tokens are released. Our experiments show that LiveMem achieves leading overall performance among evaluated systems and other intrinsic memory methods. Experiments on LongMemEval show that LiveMem is able to answer the question based on the memory state, even when the supporting evidence has been removed from the current context, and evidence-distance analysis shows that useful information persists beyond the active window. LiveMem thus establishes state continuity as a distinct and complementary abstraction for continual LLM inference.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

Decoupling Memory from Context: Structured Memory for Token-Efficient Test-Time Continual Learning

Large language models (LLMs) are increasingly deployed in enterprise, scientific, and medical applications, where agents must incorporate domain-specific knowledge and adapt from experience. Context engineering offers a practical alternative to weight updates by improving model behavior through instructions, strategies...

Yehya Farhat, Michael Desmond, Anastasios Kyrillidis · 0 citations
#artificial intelligence Preprint Sep 2026

LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory

The study of a lifecycle-labeled memory setting in which write episodes provide lifecycle metadata during training, and phase-aware readout is used during evaluation suggests that explicit lifecycle signals can help diagnose and mitigate overwrite in compact online memory.

Han-Yu Zhao, Yu-Qian Feng, Zhen-Yu Song et al. · 0 citations
#artificial intelligence Preprint Sep 2026

StateComp: Learning When to Compress History in Long Horizon Agents

This work proposes State Conditioned Compression (StateComp), a framework that determines when historical interactions can be safely compressed according to the current agent state, and demonstrates that it reduces total agent and summarization tokens while maintaining task performance, and achieves a 12.67-fold speedu...

Ming-Xuan Wang, Hong-Yue Chen, Ying-Long Guo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FlowState: Execution State as Memory for Long-Horizon LLM Agents

Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $\tau^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption.

Ming-Hao Li, Bang-Yan Li, Zi-Fan Wang et al. · 0 citations
#natural language process... Preprint Sep 2026

Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents

Memory is a core component of conversational agents, enabling coherent and context-aware behavior over long interactions. Recent approaches commonly rely on LLM-based memory construction, where raw interactions are rewritten into structured memory units and later retrieved via a RAG pipeline. While effective in control...

Dong-Hua Cai, Yong-Heng Deng, Yi-Fei Wang et al. · 0 citations
Preprint Aug 2026

MARCH: Scaling Recurrent Memory with Content-Routed State Anchors

Memory-Anchor Routing across Context History (MARCH), a network architecture that effectively scales state-space models beyond a fixed-size dimension, while maintaining computational efficiency over long-sequences, is introduced.

Ming Zhang, Kai-Sen Yang, Shu Yu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.