Large language models equipped with development environments have moved code generation toward repository-scale construction, yet building complete repositories remains difficult because interacting modules, interfaces, configurations, tests, and dependencies must work together. We introduce Code Primitives, agent-nati...
Hai-Bo Jin, Peng Kuang, Xu-Chen Yu et al.· 0 citations
Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, depende...
Hai-Bo Jin, Xin-Jie Li, Peng Kuang et al.· 0 citations
Agent harnesses govern how large language models (LLMs) gather context, invoke tools, verify results, preserve state, and terminate, largely affecting agent performance. However, the value of each harness mechanism can differ across heterogeneous tasks: a mechanism that improves one task may impose overhead or context...
Peng Kuang, Hai-Bo Jin, De-Hao Wu et al.· 0 citations
ANTMAN is introduced, an adaptive coordination framework that treats evolving unresolved information needs as the unit of runtime coordination and maintains a revisable Need Graph that tracks unresolved requirements, accumulated evidence, prior attempts, and search progress.
Jerry Wang, Hai-Bo Jin, Xiao-Peng Yuan et al.· 0 citations
BabelFlow is introduced, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-preserving translation, and combining multi-layer verification with human review to preserve task and evaluation semantics.
Peng Kuang, Yu-Chun Fan, Jiang-Nan Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.