This work evaluates three commit-time guard granularities (global epoch, read-set version, semantic commit predicate), multi-level verification, and model-side gates on three locally hosted quantized model families, and investigates how precisely runtime guards distinguish invalidating races.
Zi-Hao Zheng, Jia-Yu Long, Bai-Chuan Li et al.· 0 citations
Agentic systems commit state-changing actions, but additional verifiers can inherit the same upstream fault. We present VP-CONTROL, a runtime-assurance design and deterministic benchmark for cost-aware commit gates. Its 48 task templates yield 2,880 scenarios across six fault regimes. A fixed-call 2 x 2 experiment sepa...
Zi-Hao Zheng, Bai-Chuan Li, Jun-Yi Yao et al.· 1 citation
Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, and standard F1 scores do not indicate which predictions remain safe to automate when that input distribution changes. We study confidence estimation and selective prediction...
Zi-Hao Zheng, Bai-Chuan Li, Jun-Yi Yao et al.· 2 citations
An episode-level evaluation protocol for healthcare NLP agents is introduced, supplying a reproducible intermediate evaluation layer between static benchmarks and prospective workflow studies, with an explicit cost-sensitive treatment of missed versus unnecessary escalation.
Jun-Yi Yao, Bai-Chuan Li, Zi-Hao Zheng et al.· 1 citation
This work describes coordinated perturbation as a budgeted subset-selection problem over pairwise observations and introduces an Adaptive Subset Selection Attack (ASSA) as a scalable search heuristic for probing high-impact perturbation sets.
Junyi Yao, Zihao Zheng, Jiayu Long· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.