SKILL-KD is proposed, a contrastive skill distillation framework that treats skills as an explicit distillation medium between agents of different capabilities and consistently improves frozen student agents over fixed-model adaptation baselines.
Qi-Ming Shi, Yibo Dou, J. Zhu et al.· arXiv.org· 0 citations
This work proposes a dependency-aware code generation framework that explicitly models interactions among code entities through a graph-based representation, and introduces a sparse triplet representation for strong dependencies, significantly improving storage efficiency and computational scalability.
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumul...
Qi-Ming Shi, Yulong Tao, Linbo Jin et al.· arXiv.org· 2 citations
Computation-conditioned credit transport (CCT) is introduced, a general framework in which a detached statistic of the behavior policy's internal computation parameterizes the causal kernel that transports downstream value through a rollout.
OrderProbe is introduced, a deterministic benchmark for structural reconstruction using fixed four-character expressions in Chinese, Japanese, and Korean, which have a unique canonical order and thus support exact-match scoring.
This work presents Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget.
Eric Jiang, Zhi Zhang, Yuchen Wu et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.