KV-streams is proposed, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance, and is shown to be an efficient and lightweight plug-and-play addition to any post-training pipeline.
Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda et al.· 0 citations
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Re...
Jonathan Light, C. Cui, Jeonghye Kim et al.· 0 citations
We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through a unified interface...
Roger Creus Castanyer, Marc-Alexandre Côté, Matthew J. Sargent et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.