How much of a running workflow should a language model agent revise? ControlScope compares continuing generated code, editing the next tool call's data arguments, and replacing the unfinished workflow from the same public execution state. The nested permissions separate available repairs from the actions an agent selec...
Jing-Jie Ning, Xue-Qi Li, Yi-Bo Kong et al.· 0 citations
AI research agents must predict the effects of computational changes after budgeted experiments. WhatWorkedBench evaluates this experimental understanding through a delivered response surface of configuration scores. Agents inspect workflow code and buy measurements; exhaustive CPU references score conditional componen...
Jing-Jie Ning, Xue-Qi Li, Yi-Bo Kong et al.· 0 citations
The Discovery Certification Protocol (DCP) turns an outcome claim into an executable audit under a registered model, information boundary, and budget, and a deterministic verifier reproduces these local decisions from frozen records.
Jing-Jie Ning, Shan Zhong, Xiao-Chuan Li et al.· 1 citation
By validating decisions rather than only artifacts, this design turns adaptive search into reusable evidence wherever agents propose executable alternatives against a fixed evaluator.
The sandbox provides a search API that indexes large-scale public web corpora, namely ClueWeb22 and FineWeb, using a state-of-the-art dense retriever and approximate nearest neighbor search via DiskANN and achieves comparable latency to popular commercial APIs while ensuring stable document rankings across runs.
João Coelho, Jingjie Ning, Jingyuan He et al.· International Conference on...· 3 citations
The hidden phenomenon accuracy-blind answer churn is called and the Snapshot Compatibility Audit is introduced, which estimates excess answer churn by subtracting same-snapshot repeat disagreement from cross-snapshot disagreement.
Jing-Jie Ning, Xue-Qi Li· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.