Skip to content
Preprint

MEMONDEMAND: A Memory Management System for Large-Scale Enterprise Data

Aug 2026 · 0 citations · 41 references
Computer Science

TL;DR

On EnterpriseRAG-Bench, MEMONDEMAND outperforms the strongest published LB#1 result at every evaluated scale from 10M tokens through the complete 618M- token collection, and results on FinanceBench, HotpotQA, and FRAMES further show strong performance across financial, multi-hop, and fact-retrieval settings.

Abstract

Enterprise repositories are large, heteroge- neous, and continuously updated, making re- trieval difficult when efficient access, source- faithful evidence, and cross-query adaptation must be supported together. Enterprise mem- ory extends retrieval beyond the model con- text, but existing systems do not jointly address collection-specific hierarchy construction, low- cost routing, detailed evidence loading, and workload-aware memory updates at this scale. We introduce MEMONDEMAND, short for On- Demand Memory, a memory management sys- tem with three coordinated mechanisms: a dy- namic multi-level hierarchy that determines the abstraction structure and depth for each col- lection, dual memory at every hierarchy level that separates distilled routing from detailed evidence, and on-demand memory promotion that updates node priority under a bounded active-state budget. On EnterpriseRAG-Bench, MEMONDEMAND outperforms the strongest published LB#1 result at every evaluated scale from 10M tokens through the complete 618M- token collection, with gains of 12.23% at 10M and 4.66% at 618M. Results on FinanceBench, HotpotQA, and FRAMES further show strong performance across financial, multi-hop, and fact-retrieval settings. Together, these results establish MEMONDEMAND as an accurate, ef- ficient, and scalable memory solution for very large enterprise repositories across data scales, domains, and evidence requirements. Our code is available at https://github.com/ xfab-xinyuansong/MemOnDemand.git.

View source

Similar papers

Open access Aug 2026

Why We Created Yet Another Memory Framework: Understanding MGA's Role in Next-Gen Database Systems

The Managed Global Area is introduced, a scoped shared-memory abstraction in Oracle AI Database that allows components to explicitly define allocation source, membership, and coordination semantics across selected processes while integrating with a production database engine.

Vikramraj Sitpal, Pei-Jie Li, Shubham Kumar et al. · 0 citations
Jul 2026

MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents

The MemLens is presented, a value-aware memory management system that takes memory records as first-class data objects and can serve as an efficient, interpretable, and personalized long-term memory management system for agents.

Shuyue Wei, Chang Liu, Zi-Mu Zhou et al. · 0 citations
Open access Jul 2026

MEMTIER: Tiered Retrieval, Session-Level Injection, and Typed Consolidation for Long-Running LLM Agents

MEMTIER, a tiered memory architecture and consolidation framework for an open-source agent runtime and three questions: what to store, what to inject, and what to keep are studied and cast agent memory as a pattern recognition problem: recognizing which session patterns carry evidence and which knowledge types to retai...

Bronislav Sidik, L. Rokach · 0 citations
Preprint Aug 2026

MegaMem: A Retrieval Solution for Ultra-Large Context Windows

These results show that MegaMem supports ultra-large persistent memory while preserving strong answer accuracy under a bounded generation context, and provides a practical path toward accurate retrieval over memories ranging from hundreds of millions to one billion tokens.

Xin-Yuan Song, Bo-Wen Zhu, H. Haque et al. · 0 citations
Jul 2026

MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory

MemTxn is a governance layer outside the answer model that verifies whether an update is supported by its source and restores the application-visible state after a fault, and achieves the highest average F1 across all twelve answer-model configurations.

Han-Shuai Cui, Zhiqing Tang, Z. Yao et al. · 2 citations
Aug 2026

SOS: A High-Performance Distributed Key-Value Store for Large-Scale Online Services

Large-scale online services—including web search, recommendation, and LLM inference workloads such as Retrieval-Augmented Generation (RAG) and KV-cache offloading—demand storage that handles petabyte-scale data under millisecond tail-latency SLAs. In-memory stores are cost-prohibitive at scale; disk-based systems sacri...

Ying-Xin Li, Kai Liu, Hanglun Xie · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.