Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existing activation-manage...
Xin-Rui Chen, Mengyang Li, Ou Wu et al.· 0 citations
Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior. Existing memory methods seek to reuse past experience, but most rely on a single memory representation (e.g., trajectories, reflectio...
Yong-Xian Wei, Yi-Lin Zhao, Run-Xi Cheng et al.· 0 citations
CompKV is introduced, the first compensation-aware sparse attention framework that divides tokens into blocks and explicitly optimizes selection for the downstream compensation mechanism, and shows that the residual left by block-level mean compensation is governed by both block attention mass and within-block logit va...
Zheng-Hong Huang, Rui-Zhe Yao, Dan-Yi Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.