Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. First, information is often scattered and fragmented across the environment. Second, releva...
Yu-Kai Wu, Yuan-Jing Yang, Leon Zhou et al.· 0 citations
Graph Neural Networks (GNNs) have achieved strong performance in node classification, yet their performance often drops when facing graph distribution shifts between training and testing nodes. Existing methods have been explored to improve generalization under such shifts. However, many of them either rely on environm...
Jia-Xing Li, Jia-Shuo Liu, Wei-Huang Zheng et al.· IEEE Transactions on Pattern...· 0 citations
Env-Rethink is proposed, a system with 27B post-trained model that supports three main capabilities that adaptively builds Collection Maps and Event Logs to supplement necessary context and evolves environments through virtual event histories that alter environmental states and evidence relationships.
Yu-Kai Wu, Yuan-Jing Yang, Leon Zhou et al.· 0 citations
E-Bench is introduced, a fully synthetic benchmark with 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting, and it shows that multi-step tool use remains challenging: Pass^3 stays below 60% for the strongest models, and even with code execution in the E-Bench-Code exten...
Weihuang Zheng, Tianyuan Zou, Eileen Ye et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.