Skip to content

Author

Jun-Yi Yao

We have 8 of 10 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Aug 2026

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

This work evaluates three commit-time guard granularities (global epoch, read-set version, semantic commit predicate), multi-level verification, and model-side gates on three locally hosted quantized model families, and investigates how precisely runtime guards distinguish invalidating races.

Zi-Hao Zheng, Jia-Yu Long, Bai-Chuan Li et al. · 0 citations
Preprint Sep 2026

Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures

Agentic systems commit state-changing actions, but additional verifiers can inherit the same upstream fault. We present VP-CONTROL, a runtime-assurance design and deterministic benchmark for cost-aware commit gates. Its 48 task templates yield 2,880 scenarios across six fault regimes. A fixed-call 2 x 2 experiment sepa...

Zi-Hao Zheng, Bai-Chuan Li, Jun-Yi Yao et al. · 1 citation
Preprint Aug 2026

Reliable Financial Named Entity Recognition Under Domain Shift: Confidence Estimation and Selective Prediction

Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, and standard F1 scores do not indicate which predictions remain safe to automate when that input distribution changes. We study confidence estimation and selective prediction...

Zi-Hao Zheng, Bai-Chuan Li, Jun-Yi Yao et al. · 2 citations
#natural language process... Preprint Aug 2026

Toward Workflow-Aware Benchmarking for Healthcare NLP Agents

An episode-level evaluation protocol for healthcare NLP agents is introduced, supplying a reproducible intermediate evaluation layer between static benchmarks and prospective workflow studies, with an explicit cost-sensitive treatment of missed versus unnecessary escalation.

Jun-Yi Yao, Bai-Chuan Li, Zi-Hao Zheng et al. · 1 citation
Preprint Aug 2026

Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents

The memory-clarification boundary is studied: whether interaction-derived information should be persisted, used only in the current context, re-verified, or clarified with the user, as well as across Claude and Qwen.

Bai-Chuan Li, Jun-Yi Yao, Zi-Hao Zheng · 3 citations
#artificial intelligence Preprint Jun 2026

Beyond Helpfulness: A Teaching-over-Solving Diagnostic for Measuring Educational Impact in LLM Tutors

It is argued that public tutoring benchmarks can better support positive-impact evaluation by reporting solving-oriented and pedagogy-oriented scores separately and by making disclosure-sensitive, student-agency-preserving criteria more explicit.

Jun-Yi Yao, Zi-Hao Zheng, Baichuan Li · 3 citations
#artificial intelligence Preprint Apr 2026

Perturbation Sensitivity of Maximum-Likelihood Pairwise Ranking in Computational Decision Systems

This work describes coordinated perturbation as a budgeted subset-selection problem over pairwise observations and introduces an Adaptive Subset Selection Attack (ASSA) as a scalable search heuristic for probing high-impact perturbation sets.

Junyi Yao, Zihao Zheng, Jiayu Long · 3 citations
Conference Open access Jul 2026

JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models

JailMeter is proposed, an evidence-based evaluation framework designed to more faithfully measure jailbreak effectiveness and distill into a small language model, JailMeter\textsubscript{SLM}, which maintains comparable reliability with significantly reduced computational costs.

Qingjia Huang, Jingyu (Jack) Zhang, Jianguo Wu et al. · 2 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.