This work introduces an inclined boundary that evaluates prediction loss relative to predictive entropy, and shows that entropy correction can preserve the expected membership signal while reducing its variance, thereby improving standardized member--non-member separation.
Chen-Ye Ke, Zi-Rui Liu, Qi Liu et al.· 0 citations
CAT is introduced, a paradigm in which the agent writes Playwright code to drive the browser, gathers feedback, and autonomously explores web applications to uncover bugs, revealing a clear gap between current VLM capabilities and the demands of real-world testing in AI web development.
Bin Hong, Zhen-Chao Zhang, Ji-Yuan He et al.· 0 citations
COVE is presented, a unified agent self-evolution framework that combines harness-based and parameter-based learning through task-aware routing, stage-aware scheduling, and knowledge optimization, and shows that COVE outperforms single-channel evolution strategies.
T. Ji, Zhenya Huang, Jiayu Liu et al.· 0 citations
This paper systematically compares four state-of-the-art LLMs, including Gemini~3.1~Pro, GPT-5.4, and Claude~Opus~4, on a FormalMATH subset and on two real PDE theorems requiring deep domain expertise, evaluating their ability to produce verified Lean~4 proofs and to identify errors in deliberately incorrect proofs.
Junjie Zhang, Jia-Yin Liu, Wenbin Liu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.