Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a sca...
Xin-Ke Jiang, Tao Feng, Zhi-Bang Yang et al.· 0 citations
Long-term memory supports the self-evolution of LLM agents by retaining experience and skills across tasks and enabling their retrieval, reuse, and revision in subsequent long-horizon decision-making. Yet existing memory management approaches remain limited to discriminative retrieval and to address the sparse, hierarc...
Xin-Ke Jiang, Tao Feng, Wei-Xuan Xu et al.· 0 citations
Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is reflection: assessing trajectory progress, identifying missing evidence and unreliable intermediate states, and deciding whether to continue,...
Zhi-Xin Zhang, Xin-Ke Jiang, Zhi-Bang Yang et al.· 0 citations
Harness-RL is introduced, a structured reinforcement learning framework that combines Conflict-Aware Policy Optimization (CAPO) with interface-level black-box trajectory construction and supports both central-only and joint multi-agent training.
Xin-Ke Jiang, Zhi-Xin Zhang, Zhi-Bang Yang et al.· 1 citation
ScaffoldAgent is proposed, a utility-guided dynamic outline optimization framework for OEDR that models outline evolution as a structured decision process with three operations: Expansion, Contraction, and Revision, enabling controlled updates to the report scaffold.
This work proposes AgenticRag-R1, a RL framework that deeply integrates reasoning, retrieval, and memory via a memory stack and fine-grained action space, supported by hierarchical action-aware rewards and an information-aware trajectory rejection strategy to enable effective long-horizon learning.
Xin-Ke Jiang, Yue Fang, Zhi-Bang Yang et al.· 2 citations
CoHarden is proposed, a co-generation framework that uses the Lax signal as an in-loop convergence criterion that generates a test before any fix, then iteratively hardens the test and fix against surviving mutation patches until the generated test no longer admits Lax regressions.
Yuhao Tan, Zhibang Yang, Fangkai Yang et al.· arXiv.org· 1 citation
This work introduces ToolAtlas, a graph-based framework that builds a persistent provider-side tool memory of tool capabilities, failure boundaries, and cross-tool compositions through execution-verified probing and establishes provider-side tool memory as an effective and reusable paradigm for tool servers.
Yue Fang, Zhibang Yang, Fangkai Yang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.