Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At c...
Yun-Tian Deng, Peng-Yu Nie, Stuart M. Shieber· 0 citations
It is found that a good bug report for an agent overlaps with, but is not identical to, a good report for a human: agents benefit most from concrete, executable, and well-localized information, whereas some qualities long emphasized for human readers, such as natural language steps to reproduce and readable description...
Lara Khatib, N. Mathews, M. Nagappan et al.· arXiv.org· 2 citations· ⚡1
TestEvo-Bench is a benchmark of test and code co-evolution tasks mined from software repositories, with two tracks: in test generation, the agent shall write new tests to capture the new software behavior; in test update, the agent shall adapt failing existing tests to the changed software behavior.
It is confirmed that memorization persists in modern LLMs and is influenced more by a complex interplay of training domain, dataset composition, architectural choices, and content characteristics, rather than parameter count alone.
Xiaoyu Cheng, Kundi Yao, Pengyu Nie et al.· AIware· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.