Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee com...
Bo Mao, Hang He, Lin-Ting Wang et al.· 0 citations
ContextWeave is introduced, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams and motivates memory systems that optimize not only retrieval relevance but also reliable use during execution.
It is found that memory can improve UAQ performance in some settings, but such gains are selective rather than universal and remain fragile under dataset shift, suggesting that reliable UAQ memory depends less on storing larger amounts of experience and more on preserving transferable behavioral guidance.
Chuanyuan Tan, Junpu Yu, Yuxia Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.