ReMEMBER is proposed, a missing-evidence memory framework that conditions retrieval on unresolved window dependencies and refines retrieved chunks into evidence-dense memory under a fixed budget and improves memory recall and gap-resolution completeness over memory construction baselines under the same budget.
Abstract
Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on its own. We formalize this setting as streaming dialogue summarization, where a system must summarize a current window using selective memory from an unbounded history under a fixed budget. We show that the central challenge is not how much history is accessed, but whether memory recovers the evidence that the current window presupposes. We construct a benchmark and evaluation protocol that separately assesses whether memory contains gap-resolving evidence and whether the generated summary reflects it. We propose ReMEMBER, a missing-evidence memory framework that conditions retrieval on unresolved window dependencies and refines retrieved chunks into evidence-dense memory under a fixed budget. Experiments on dialogues with histories up to 160K tokens show that ReMEMBER improves memory recall and gap-resolution completeness over memory construction baselines under the same budget.
REMAP is proposed, a reflection-guided memory editing approach for online alignment of persona facts that selectively writes and revises memory entries based on the current dialogue evidence and retrieved related items, enabling more efficient context utilization over extended interaction horizons.
Qingyang Xu, Xiao Liu, Zhou Fang et al.· Annual International ACM SIG...· 0 citations
Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We introduce LOCOMO-CONV, a conversa- tional memory b...
This work investigates whether memory interference originates mainly from memory retrieval or from the accumulation of competing fact versions added during memory updates, and evaluates how memory-write policies influence memory retrieval behavior later on.
Erica Butts, Salam Daher· Proceedings of the 26th ACM...· 0 citations
RUMBA (Russian User Memory BenchmArk) is introduced - a new benchmark for long-term conversational memory that provides a fine-grained taxonomy of memory-centric question types and a unified methodology accounting for semantic type, session scope, temporal reasoning, and the explicitness of temporal expressions.
E.D. Shevtsova, Inna Glebkina, Mark Baushenko et al.· arXiv.org· 0 citations
It is hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response).
Ryuichi Sumida, K. Inoue, Tatsuya Kawahara· 0 citations
The results indicate that for precise, evidence-grounded questions over chat archives, much of the benefit credited to elaborate memory structures is recoverable by giving an agent controllable search over the unmodified record, with no LLM-based index construction at all.
Ruizhe Li, L. Zhang, Benfeng Xu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.