Skip to content

Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating

Jul 2026 · arXiv.org · Vol abs/2607.24667 · 0 citations · 19 references
Computer Science

Abstract

A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, whether an item will be reused, placing existing methods on one axis, the commit lag $H$: online filters and learned predictors commit at $H=0$, while Belady's offline optimum sits where the whole future is known. The missing regime in between, fixed-lag smoothing, waits a bounded number of steps, observes which items a correct near-future prediction attended to, and only then commits. This measurement, demonstrated utility, turns Belady's unobservable future request into something we read off the model itself. We instantiate it as a training-free policy, RMM, a strict generalization of H2O that reduces to it exactly when the measurement is uniform. In controlled settings where reuse is endogenous and separated in time, demonstrated utility identifies used memory far better than accumulated attention, and a small bounded memory behaves like a much larger one. But on independent third-party benchmarks, run inside NVIDIA's KVPress harness against its own SnapKV, H2O, and StreamingLLM implementations, the advantage mostly disappears: RMM is on par with H2O for single-turn question answering and loses to both H2O and SnapKV in a streaming multi-turn setting. The cause is simple: on natural text the model is correct about most tokens, so weighting attention by correctness barely changes it, and demonstrated utility collapses onto accumulated attention unless reuse is sharp and endogenous, which standard benchmarks do not exercise. Our contribution is the framework and an honest map of when measuring beats accumulating, not a new state of the art.

View source

Similar papers

Preprint Aug 2026

Evidence Before Expansion: Reuse, Spawn, or Defer in Lifelong Expert Pools

This work proves finite-time anytime validity for the observable surrogate discrepancy of a predictable discriminator sequence, and an unconditional one-sided transfer to the population quantity in which each side's slack is the excess risk of a single discriminator.

Kentaro Oda · 0 citations
#natural language process... Preprint Aug 2026

Hindsight Memory-PRM: Supervising Memory Management with Auditable Hindsight Credit

Hindsight Memory-PRM exploits this audit trail twice: offline to train an operation-conditioned memory-utility critic, and online, where retrievals, citations, and one controlled deletion-and-reanswer per probe settle an intervention-calibrated entry-level presence credit.

H. Jia, Yang Liu, Ying-Guang Yang et al. · 0 citations
Preprint Sep 2026

A Historical Corpus Is Not a Historical System: Auditing Hindsight Leakage in Stateful Data Discovery

Offline replay should estimate what a discovery system could retrieve at a historical point, yet freezing the corpus leaves interaction memory unconstrained. We formalize point-in-time (PIT) discovery through historical state $(D_t, \theta_t, M_{<i})$ and introduce a paired replay that changes only memory availability....

Yi-Xi Zhou, Fan Zhang, Si-Kun Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation

The fork ledger is introduced, which branches a deployment stream at pre-registered decision points into matched update and hold continuations under common random numbers, allowing triggers to be judged by the updates they select rather than by surprise detection alone.

An-Qi Li, Kaden Kim · 0 citations
Preprint Sep 2026

When is Test-Time Adaptation Identifiable From Unlabeled Evidence?

The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.

Kartik Jhawar, Li-Po Wang · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.