Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this problem, we introduce the Correlated Promotion Benchmark (CP...
Xiao-Yang Li, Yi-Qi Wang, Chen-Cheng Zhu et al.· 0 citations
This work proposes BiVCoder, a diagnosis-driven multi-agent framework featuring a novel bidirectional code-test diagnosis mechanism, and introduces BiVCoder-SFT, a role-specific instruction fine-tuning scheme.
Xiaoyang Li, Jin-Hao Dong, Wenhang Shi et al.· Proceedings of the 32nd ACM...· 0 citations
This work argues that a memory write is not a belief commit, and presents MemTX, a transactional belief-commit protocol, a transactional belief-commit protocol that leads all eight baselines with paired-McNemar significance on four backbones and statistically ties the best baseline on the fifth and strongest, while rem...
Xiaoyang Li, Yi-Qi Wang, Haohui Lu et al.· arXiv.org· 5 citations· ⚡1
Large Language Models (LLMs) have demonstrated remarkable potential in automated code generation. However, existing test-driven code generation and refinement frameworks are often hindered by the tests' quality: they typically treat self-generated tests as ground truth, leading to ineffective debugging loops where code...
Xiaoyang Li, Jinhao Dong, Wenhang Shi et al.· Proceedings of the 32nd ACM...· 0 citations
Weight-space composition supports coarse, input- and format-conditioned functional statements -- not a universal merging-performance predictor, and not one that training-format evaluations can see.
A one-round study provides initial evidence for PRD-guided self-evolution, motivating validation at larger scales and in industrial settings, and presents AgentOmnia, a framework coordinating task-space definition, data synthesis, post-training, evaluation, and improvement across To-Consumer (ToC), To-Business (ToB), a...