Skip to content

RUMBA: Russian User Memory Benchmark

Jul 2026 · arXiv.org · Vol abs/2607.21447 · 0 citations · 5 references
Computer Science

TL;DR

RUMBA (Russian User Memory BenchmArk) is introduced - a new benchmark for long-term conversational memory that provides a fine-grained taxonomy of memory-centric question types and a unified methodology accounting for semantic type, session scope, temporal reasoning, and the explicitness of temporal expressions.

Abstract

The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate retrieval metrics, failing to capture interactions between long-range context, temporal information, and reasoning. To address this, we introduce RUMBA (Russian User Memory BenchmArk) - a new benchmark for long-term conversational memory that provides a fine-grained taxonomy of memory-centric question types and a unified methodology accounting for semantic type, session scope, temporal reasoning, and the explicitness of temporal expressions. RUMBA consists of timestamped user-assistant dialogues with QA pairs requiring retrieval, combination, and reasoning across sessions. While designed for Russian, we also provide an aligned English subset under the same methodology. We evaluate contemporary memory systems and long-context models, and show how RUMBA serves as a diagnostic tool to analyze model behavior across benchmark slices and identify strengths and failure modes of different memory mechanisms.

View source

Similar papers

#natural language process... Preprint Aug 2026

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UtilMem is introduced, a diagnostic benchmark comprising 1,717 instances across five domains, designed to evaluate four underexplored aspects of memory utilization: reasoning over dense histories, identifying implicitly relevant memories, synthesizing distributed evidence into summaries, analyses, or plans, and resisti...

Pei-Jun Qing, Fobo Shi, S. Vosoughi · 0 citations
#artificial intelligence Preprint Sep 2026

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We introduce LOCOMO-CONV, a conversa- tional memory b...

Wen-Yu Chang, Yun-Nung Chen · 3 citations · ⚡1
#artificial intelligence Preprint Jul 2026

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

A benchmark and harness that measures the marginal utility of memory for task-executing agents under explicit cost accounting, and provides episodic tool-use tasks in three domains whose dependence on earlier-episode facts is verified by an automated leak check.

Shweta Mishra, Shashank Mishra · 0 citations
Aug 2026

DynaGraph-LLM: a dynamic ontological memory framework with multi-scale retrieval for mitigating contextual amnesia in large language models

DynaGraph-LLM is introduced, a novel neuro-symbolic architecture that endows LLMs with a dynamic, persistent, and structured memory and implements a Dual-Phase Memory Consolidation process, inspired by hippocampal-neocortical interactions in the human brain, to refine and abstract knowledge over time.

Abdelweheb Gueddes, B. Louhichi, Mohamed Ali Mahjoub · 0 citations
Preprint Aug 2026

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

It is hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response).

Ryuichi Sumida, K. Inoue, Tatsuya Kawahara · 0 citations
#natural language process... Preprint Sep 2026

Where to Look and What to Use: Retrieve-Localize-Generate for Long-Term Conversational Memory Question Answering

This work proposes MemLoc, a unified Retrieve-Localize-Generate framework for long-term conversational memory QA, and introduces a reasoning-based evidence locator trained with Self-reflective Hint Policy Optimization, which performs progressive refinement by extracting query-relevant fragments within memory units to s...

Yi-Fan Wang, Xin-Kui Lin, Yong-Xiu Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.