Skip to content
Preprint

HyMem: Hierarchical Context Management for Long-Horizon Agents via Information Isolation

Aug 2026 · 0 citations · 41 references
Computer Science

TL;DR

HyMem is a hierarchical framework that explicitly separates the agent's context into distinct functional layers to separate high-level planning from execution and complex analysis, allowing the model to maintain focus and accuracy across complex, long-horizon tasks.

Abstract

Large language model (LLM) agents often perform poorly on complex, long-horizon tasks because their context becomes increasingly cluttered over time. As interactions accumulate, detailed execution traces and intermediate outputs dominate the context, making it difficult for the model to retain and use high-level planning information. Most existing methods address this issue through compression or retrieval applied to a single, flat context, which does not clearly separate different types of context information and often leads to degraded reasoning. To address this challenge, we propose HyMem, a hierarchical framework that explicitly separates the agent's context into distinct functional layers. HyMem organizes context by function to separate high-level planning from execution and complex analysis. Its isolated reasoning module handles complex subtasks without adding intermediate reasoning traces to the persistent planning context, while its memory management module preserves task progress across context refreshes through structured summaries. These components reduce redundant context accumulation, retain task-critical information, and support coherent long-horizon reasoning within a limited context window. Experiments on GAIA and Browsecomp-plus show that, with DeepSeek-V4, HyMem achieves average Pass@1 scores of 66.7% and 61.3%, outperforming the strongest baseline by 6.1 and 4.7 percentage points, respectively. Further analysis indicates that HyMem effectively controls the growth of the reasoning context, allowing the model to maintain focus and accuracy across complex, long-horizon tasks.

View source

Similar papers

Jul 2026

ACM: Agentic Context Management for Long Horizon Tasks

This work proposes Agentic Context Management (ACM), a framework that equips agents with purpose-built context editing tools for lossless context management and develops a post-training pipeline that constructs high-quality demonstrations of context management and improves model performance on both agentic search and c...

Xiao-Chuan Li, Ryan Ming, Meng Chu et al. · 1 citation
#natural language process... Preprint Aug 2026

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

An RL method tailored for context management is proposed, which uses context and entropy variation to identify critical editing decisions for branch sampling and estimates action-level advantages from all branched trajectories that pass through the corresponding context editing action.

Zhuo-Shi Pan, Qizhi Pei, Jun-Ru Lu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems

Gated-Memory Routing is proposed, which conditions each decision on the query and a learned execution memory, and attains the best average accuracy, exceeding the strongest baseline by 2.44 points, while reducing HumanEval inference cost by 31.9% relative to that baseline.

Rakibul Hasan Rajib, Meng Zheng, Qian Lou · 1 citation
Preprint Aug 2026

Context as an Environment: Programmatic Context Management for Long-Horizon Agents

LLM agents increasingly take on long-running tasks whose history grows far beyond a single model context window. Existing approaches compress earlier interactions or extract selected information into fixed memory representations, committing to what to preserve before future needs are known. We present Scroll, a context...

Yin Lin, Elaine Ang, E. Zhu et al. · 2 citations
Jul 2026

PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning

Pro-LONG, a minimal context management framework built around programmatic memory for LLM agents in long-horizon, exploratory settings, is proposed, which addresses the tradeoff of preserving more information makes retrieving relevant details less tractable.

A. Fox, Jun-Lin Wang, P. Rosu et al. · 2 citations · ⚡1
Preprint Aug 2026

Chained Recursive Language Models for Multi-Iteration Reasoning

This work proposes Chained Recursive Language Models (Chained RLM), an inference-time architecture, in which the same underlying model is called repeatedly as a sequence of fresh reasoning roots, and studies when fresh-context artifact continuation gives a measurable gain in accuracy over direct LLM answering even with...

Purbesh Mitra, S. Ulukus · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.