Skip to content

Author

Ya-Liang Li

We have 6 of 85 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics, Convergence Rates, and Benefits of Off-Policyness

A unified theory for RE(S) is developed that covers the full spectrum of S, and can be interpreted as a stage-wise optimization process, where each stage takes $S$ gradient steps for minimizing the Kullback-Leibler distance to a fixed reward-weighted rollout distribution.

Zhi-Wei Wang, Yan-Xi Chen, Ya-Liang Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Speculative Evaluation of Stochastic LLMs

Real-generation experiments that account for the pilot synchronization barrier show that HBN-async mitigates its overhead, helping translate statistical efficiency into practical evaluation benefits, and outperforming hindsight-tuned empirical and independent Bayesian baselines.

Qian-Li Shen, Xiang Li, Ruo-Meng Ding et al. · 0 citations
Preprint Aug 2026

Private Direct Preference Optimization for LLM Alignment

This paper formalizes preference privacy, a label-DP-style privacy notion for DPO that protects only the relative preference between candidate responses, assuming an adversary who already knows the prompt and responses, and designs PrivDPO, a DPO variant that enforces preference privacy while remaining compatible with...

Yang-Fan Jiang, Fei Wei, Ergute Bao et al. · 0 citations
Preprint Aug 2026

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

TrajVal, a lightweight probe-based estimator that approximates per-task learnability from a short probe run and two endpoint evaluations, is proposed and it is found that learnability is reproducible across independently sampled training contexts and predictive of downstream utility.

Ting Zhou, Zhenqing Ling, Daoyuan Chen et al. · 0 citations
Jul 2026

Hybrid Analysis for Secure MCP Tool Use in LLM Agents

MTGuard is proposed, a hybrid analysis-based defense framework designed to safeguard the use of MCP tools in LLM agents by leveraging lifecycle-aware static-dynamic co-analysis and effectively mitigates multiple categories of harmful tool use across different LLM agents while maintaining performance on benign user task...

Ping He, Yuexiang Xie, Yaliang Li et al. · 0 citations
Preprint Jul 2026

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

EvoSOP is introduced, a framework that empowers agents to extract SOPs from execution trajectories and iteratively optimize the toolset through a systematic lifecycle of construction, merging, evaluation, and pruning, providing a scalable pathway for the development of self-evolving agents.

Haipeng Ding, Yuexiang Xie, Zhewei Wei et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.