Skip to content
Preprint

CrystalMem: Elastic Memory for Self-Evolving LLM Agents via Knowledge Crystallization

Jul 2026 · 3 citations · 72 references
Computer Science

TL;DR

CrystalMem (Crystallized Memory), an elastic memory sidecar that demotes entries across four fidelity states under a crystallization-energy schedule, orders demotions by advantage-weighted influence with dependency coupling, and recovers capability through verified recrystallization under explicit compute and byte caps is proposed.

Abstract

Memory for self-evolving large language model (LLM) agents is often provisioned as if its byte budget only grows. Cloud platforms, however, adjust quotas with load and cost, and we show that capability does not follow the budget back up: after a squeeze-and-recover cycle, the agent settles below its pre-squeeze level, a gap we call memory hysteresis. The cause is structural. Deletion and one-way compression discard the material needed for later rebuilding, and we prove that any policy that only keeps or drops entries carries a residual-deficit floor. We propose CrystalMem (Crystallized Memory), an elastic memory sidecar that demotes entries across four fidelity states under a crystallization-energy schedule, orders demotions by advantage-weighted influence with dependency coupling, and recovers capability through verified recrystallization under explicit compute and byte caps. Across seven environments, seventeen methods, and six backbones, with multi-tenant serving and a physical edge-cloud deployment, CrystalMem achieves the highest restored capability in every setting and closes the loop left open by every baseline. From a 50% byte budget, CrystalMem matches the strongest budgeted baseline at full provision on every environment; at equal budgets, it leads by +4.6 pp on average.

View source

Similar papers

Preprint Aug 2026

Runtime Observability for Heterogeneous Attention Memory

A runtime observability contract is given that covers all four memory classes with three operators, instantiate it on six model configurations across five architecture families, and compose the per-stage bounds into an executable request-level risk ledger.

Fanzhe Wei, Li Liu, Ziyang Wang et al. · 2 citations
Preprint Aug 2026

SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents

We present SuperLocalMemory 4.0, a governed, local-first memory operating system for AI agents, unifying multi-channel retrieval under reciprocal-rank fusion, bi-temporal recall, multi-scope isolation, role-based access, verified erasure, and a hash-chained audit trail. A reliability spine governs the primary write pat...

V. Bhardwaj, Garima Singh, Arun Pratap Bhardwaj · 1 citation
#artificial intelligence Preprint Sep 2026

Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents

Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memory, pending tool plans, and, under every serving API, a KV cache. Yet today's"forget"operations delete a plaintext memory record and stop, leaving every artifact derived from the revoked information intact. We f...

Chao Yao, Yangbo Wei, Zhen Huang et al. · 1 citation
#natural language process... Preprint Aug 2026

Dual-Layer Agentic Memory with Fast Write Routing and Slow Consolidation

Dual-Layer Agentic Memory is proposed, a framework that shifts memory management to the write phase through cost-aware epistemic routing and periodic parametric consolidation, allowing the router to adaptively suppress redundant writes as the model's epistemic boundaries evolve.

Wenzhi Li, Dong Nie, Ruiyi Lan et al. · 0 citations
Preprint Sep 2026

UNISON: A Co-Designed Near-Memory Scheduler of Session KV Residency for LLM Agents

Large language models are increasingly composed into agent loops that plan, call tools, and resume the same task after each action. These loops press a shared memory hierarchy harder than conventional multi-turn chat, because they hold a growing key-value (KV) prefix across tool waits and place many sessions on one SRA...

Fan He, Yan Li, Xiao-Yang Zeng · 0 citations
Preprint Sep 2026

Fast Recovery for LLM Serving via Decoupled Device Memory Lifetime in Dynamo

Large language model (LLM) inference replicas run across tightly coupled GPUs and serve traffic continuously for weeks. Hardware and software failures are therefore inevitable, and one worker failure can disrupt an entire replica. Recovery requires reinitializing the engine, taking minutes even when weights and compila...

Schwinn Saereesitthipitak, Mohammed H. Abdulwahhab, Hannah Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.