Skip to content

PReM: Learning What to Preserve and When to Refresh for Context Compression

Jul 2026 · arXiv.org · Vol abs/2607.14327 · 0 citations · 27 references
Computer Science

TL;DR

PReM (Preserve and Refresh Memory), a context-compression framework that maintains the long context as the model's internal layer-wise KV memory and learns what to preserve and when to refresh it, is introduced.

Abstract

Efficient long-context inference is not only about reducing memory cost, but also about keeping useful contextual evidence accessible as generation proceeds. However, existing compression-oriented approaches, such as key-value (KV) cache compression and context compression, often either make an early decision about which contextual information to keep or rely on an external compressor. Such designs make it difficult to adapt the compressed context to the evidence needed by later reasoning steps. This paper introduces PReM (Preserve and Refresh Memory), a context-compression framework that maintains the long context as the model's internal layer-wise KV memory and learns what to preserve and when to refresh it. Specifically, PReM uses a dedicated memory layer to make memory-selection decisions, and a special memory tokento trigger refreshes during generation. To train this behavior, PReM introduces Phase-Separated Refresh Training, aligning memory selection with memory-conditioned generation while preserving continuity across refreshes. Experiments with 32K-token contexts show that PReM outperforms strong baselines under both 16x and 32x compression, while maintaining a favorable balance between answer quality and inference efficiency.

View source

Similar papers

Preprint Aug 2026

From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG

Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overhead: retrieved context lengthens the prompt, increasing prefill work, KV-cache footprint, memory traffic, latency, and energy. Context compression offers a natural remedy by pruning...

Z. Feric, Amir Taherin, Yanzhi Wang et al. · 0 citations
Book Open access Jul 2026

C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference

C2KV is proposed, a unified framework for non-prefix KV reuse that jointly optimizes KV cache compression and concatenation that significantly reduces KV cache storage and transfer costs.

Chuheng Du, Jun-Yi Chen, Hanlin Tang et al. · 2 citations
#machine learning Preprint Aug 2026

Compression-Aware Abstention: Teaching LLMs to Refuse When KV-Compression Masks Remove Answer Evidence

This is the first work to formulate compression-aware abstention as a learning problem, in which a model learns to answer when supporting evidence survives compression and abstain when it does not, and controlled-deletion experiments show that the learned behavior is driven by evidence content rather than input length...

Mohammadali Khodabandehlou, Bhaskar Krishnamachari · 0 citations
Preprint Aug 2026

Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

This work instantiates a new paradigm of incremental memory activation, where the effective capacity of memory is progressively expanded as the context grows, and applies this paradigm to state-of-the-art models, observing consistent improvements on standard language modeling and reasoning, as well as on long-context r...

Reza Bayat, Ali Behrouz, V. Mirrokni et al. · 0 citations
#natural language process... Preprint Aug 2026

When to Adapt: Conditional Memory Adapters for Retention-Preserving Domain Specialization

Engram Adapter, a framework that repurposes pretraining-time conditional memory as a post-hoc adapter for frozen LLMs, improves in-domain accuracy while preserving 99.4%--100.1% of average OOD performance; on LegalBench it slightly exceeds the frozen base model on average, whereas comparable always-on baselines degrade...

Jiaxuan Hou, Lei Wang · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.