This work evaluates three commit-time guard granularities (global epoch, read-set version, semantic commit predicate), multi-level verification, and model-side gates on three locally hosted quantized model families, and investigates how precisely runtime guards distinguish invalidating races.
Zi-Hao Zheng, Jia-Yu Long, Bai-Chuan Li et al.· 0 citations
Agentic systems commit state-changing actions, but additional verifiers can inherit the same upstream fault. We present VP-CONTROL, a runtime-assurance design and deterministic benchmark for cost-aware commit gates. Its 48 task templates yield 2,880 scenarios across six fault regimes. A fixed-call 2 x 2 experiment sepa...
Zi-Hao Zheng, Bai-Chuan Li, Jun-Yi Yao et al.· 1 citation
Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, and standard F1 scores do not indicate which predictions remain safe to automate when that input distribution changes. We study confidence estimation and selective prediction...
Zi-Hao Zheng, Bai-Chuan Li, Jun-Yi Yao et al.· 2 citations
An episode-level evaluation protocol for healthcare NLP agents is introduced, supplying a reproducible intermediate evaluation layer between static benchmarks and prospective workflow studies, with an explicit cost-sensitive treatment of missed versus unnecessary escalation.
Jun-Yi Yao, Bai-Chuan Li, Zi-Hao Zheng et al.· 1 citation
The memory-clarification boundary is studied: whether interaction-derived information should be persisted, used only in the current context, re-verified, or clarified with the user, as well as across Claude and Qwen.
It is argued that public tutoring benchmarks can better support positive-impact evaluation by reporting solving-oriented and pedagogy-oriented scores separately and by making disclosure-sensitive, student-agency-preserving criteria more explicit.