Skip to content

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

Sep 2026 · 0 citations · 41 references
Computer Science

TL;DR

MemCalib-RL is proposed, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation and achieves the best overall performance while better balancing over-use and under-use.

Abstract

The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative policy optimization and on-policy self-distillation, further reveal a clear directional skew: trained models improve in one direction while deteriorating in the other. We therefore propose MemCalib-RL, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation. Results across model families and scales (Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B) show that MemCalib-RL achieves the best overall performance while better balancing over-use and under-use, with gains generalizing beyond MemCalib in external benchmark evaluation. Further experiments support its design choices and robustness and provide insight into its training dynamics.

View source

Similar papers

Preprint Aug 2026

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

This work identifies memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance, and proposes AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps.

Mengru Wang, Haozhe Luo, Zhen-Qiang Xu et al. · 0 citations
Preprint Aug 2026

Understanding Stage-Wise Utility-Risk Trade-offs in LLM Agent Memory

Control evaluations reveal that targeted poisoning risk varies across memory operations and motivate stage-aware evaluation and control of LLM-agent memory, showing that targeted poisoning risk varies across memory operations.

Chuan-Chao Zang, Zi-Jian Cao, Xiang-Tao Meng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents

Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior. Existing memory methods seek to reuse past experience, but most rely on a single memory representation (e.g., trajectories, reflectio...

Yong-Xian Wei, Yi-Lin Zhao, Run-Xi Cheng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DolphinBench: Mapping the Pareto Frontier of Agent Memory

DolphinBench is presented, a benchmark that evaluates memory directly through an agent's task completion and requires all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically.

Soumil Rathi, Deshraj Yadav, Taranjeet Singh · 0 citations
Preprint Aug 2026

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

A controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms, shows that no single substrate consistently dominates.

Wei-Chieh Huang, Wei-Zhi Zhang, Yu-Chen Wu et al. · 2 citations
#machine learning Preprint Sep 2026

Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs

In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction tuning changes how models behave under conflict, but whether it rewires the underlying circuit or merely gates/reweights already present components, remains unknown. We...

S. Pandere, Gautam Ranka, Ritika Varshney et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.