Skip to content
Preprint

Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents

Aug 2026 · 3 citations · 18 references
Computer Science

TL;DR

The memory-clarification boundary is studied: whether interaction-derived information should be persisted, used only in the current context, re-verified, or clarified with the user, as well as across Claude and Qwen.

Abstract

Persistent memory can personalize an LLM agent, but an incorrect durable update can silently distort future behavior. We study the memory-clarification boundary: whether interaction-derived information should be persisted, used only in the current context, re-verified, or clarified with the user. MCB contains 140 primary scenarios, split into 70 development and 70 held-out items, plus a separate 70-item contrast set. It evaluates both action labels and structured tool-call selection. Two non-authors independently label the 70 held-out primary and 70 contrast items (97.1% agreement, Cohen's kappa = 0.962); a blind third resolves four disagreements, replacing eight author labels by non-author majority. Across Claude and Qwen, models verify changing facts more reliably than they ask users to resolve ambiguity. Bare Qwen asks on 0/12 clarification items while verifying 12/18 freshness items. Few-shot prompting raises accuracy from 0.557 to 0.771 (paired delta = +0.214, Holm-adjusted exact McNemar p_H = 0.002), yet clarification recall remains 0.333. The policy prompt reduces erroneous persistence from 0.243 to 0.100 (p_H = 0.038), although its accuracy gain is not significant. Label-tool agreement is 57% for each Claude model and 23% for Qwen; Qwen accuracy falls from 0.557 to 0.343 (p_H = 0.047). Memory evaluation must test both stated decisions and tool-call choices.

View source

Similar papers

#artificial intelligence Conference Open access Sep 2026

Self-Reports Are Not Verification

An environment-grounded audit is introduced in which every intermediate proposal receives an exact outcome in an evolutionary Contexto search whose feedback function assigns every valid guess an exact rank without human annotation.

En-Rong Pan, Ryan Zhou, Ting Hu · 0 citations
#artificial intelligence Preprint Sep 2026

Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents

Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a versioned, update-governed identity contract. It separates recall, composition, behavioral enactment, resistance, persistence, line...

Zhen-Yu Zhao, Roy Zhao · 0 citations
Preprint Sep 2026

ToMAS: A Pilot Failure-Grounded Theory-of-Mind Benchmark from Multi-Agent LLM Failures

LLM-based multi-agent systems can fail even when communication succeeds because agents do not correctly track their peers'roles, knowledge, or intentions. We investigate whether such inter-agent misalignment cases, labelled FC2 in MAST-Data, can be converted into functional partner-state reasoning items. ToMAS applies...

M. Ishfaq, Glaucia Melo · 0 citations
Conference Aug 2026

VAIL: Retrieval Authorizes Inspection, Not Use — Governed Memory for Long-Horizon Multi-Agent LLM Systems

As language agents persist and collaborate over long horizons, a stored fact is no longer disposable context: once recalled, it steers tool calls, planning, and cross-agent agreement. We argue that memory reliability breaks down into two distinct failure modes: staleness (the world changed after a fact was stored) and...

Ming Wang, Ke-Yang Han, Ru-Yi Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Available but Unclaimed: An Empirical Study of Human-AI Synergy

People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-L...

Robin Welsch, Michelle Rausch, Pascal Knierim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.