An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable...
De-Hai Min, Dao-An Zhang, Yiming Zeng et al.· 0 citations
Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing claims and invoking external search. Treating claims independently makes LLM and search calls scale with claim count and causes repeated searches for overlapping evidence a...
Ke-Ning Zheng, Ao-Ying Zheng, Zhi-Gang Chang et al.· 0 citations
Dr. Claw is presented, an open-source workspace that wraps existing coding-agent executors in a controllable and auditable human-in-the-loop workflow rather than introducing another autonomous agent.
D. Song, Han-Rong Zhang, Dawei Liu et al.· 0 citations
This work instantiates AutoCRAT, a decoder-side controller for frozen backbones that operates over a discrete action space and updates control decisions only at semantic boundaries, improving stability while remaining responsive to the evolving reasoning process.
Han-Jun Luo, Qiu-Shi Liu, Jing-Yang Zhang et al.· 0 citations
A controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms, shows that no single substrate consistently dominates.
Wei-Chieh Huang, Wei-Zhi Zhang, Yu-Chen Wu et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.