Large language models (LLMs) are increasingly used in software pipelines, raising concerns about harmful behaviors in security-critical domains. Existing safety evaluations predominantly probe models with single prompts or short interactions, and therefore do not capture how safety behaves under multi-step workflows wh...
Lu Yan, Zhuo Zhang, Xiang-Zhe Xu et al.· Proceedings of the ACM on So...· 0 citations
Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs...
Hao-Min Qi, Xiang-Zhe Xu, Yi-Ming Huang et al.· 0 citations
Artic is proposed, an artifact-driven workflow compiler that transforms a natural-language workflow into an artifact-driven workflow in which each step declares the artifacts it reads and writes, constraints gate produced artifacts, and explicit control transfers route execution.
Xiang-Zhe Xu, Hanxi Guo, Guangyu Shen et al.· 0 citations
This work compares six tool architectures that hold the underlying information and actions similar while varying how they are organized and exposed to the model, across three actors and a total of 11,700 trajectories to show that, even when tools provide similar capabilities, tool architecture changes agent behavior.
Xiangzhe Xu, H. Saghir, Qian-Hui Wu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.