Skip to content

Author

Wei Lin

We have 6 of 47 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning

Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already present in the base model. KL regularization is widely used to mitigate such forgetting by constraining policy drift toward a reference model. However, standard full-policy KL regularization constrains the entire response distribution and may unnecessarily restrict exploration and target-task learning. This raises a natural question: can a more precise constraint preserve existing capabilities while minimizing interference with learning new tasks? To this end, we propose \underline{Co}rrectness-Conditioned \underline{KL} Regularization (CoKL), a conditional regularization framework that narrows the preservation constraint from the full output distribution to correctness-conditioned response distributions. We instantiate CoKL with forward KL divergence and derive a practical finite-group training objective for RL-based LLM post-training. At the population level, CoKL decouples the total probability assigned to correct responses from their correctness-conditioned distribution, thereby regularizing the relative probability allocation among reference-supported correct responses without directly anchoring incorrect outputs or total correctness mass. We further show that full-policy forward and reverse KL regularization induce a strict optimal correctness gap when the reference policy is imperfect, whereas CoKL avoids this limitation. Experiments in controlled multi-solution environments and continual post-training settings across multiple model scales demonstrate that CoKL achieves a more favorable balance between target-task improvement and prior-capability retention than existing regularization methods. Our code is available at https://github.com/Lumina04/CoKL.

Li Wang, Xiaodong Lu, Xiaohan Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents

This work proposes ATLAS, a dual-horizon diagnostic evaluation framework for industrial tool-use agents that instantiates LLM judge interfaces as executable signals with explicit evidence scopes and decision boundaries, and evaluates ATLAS on Meituan Xiaotuan production traffic.

Wei Chen, Pei-Lun Zhou, Zhao-Yu Hu et al. · 0 citations
Preprint Aug 2026

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

A new task, Behavior-Aware Travel Planning, which infers user preferences directly from past behaviors and generates personalized travel plans and proposes B2T-Agent, a reinforcement learning-based agent that leverages user behavior trajectories, interacts with external tools for preference-aligned retrieval, and maintains an internal memory module.

Zihao Cheng, Yingyu Shan, Hongru Wang et al. · 0 citations
Preprint Aug 2026

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

HiDiffTIR is proposed, a Hierarchical Difficulty-aware policy optimization framework for multi-turn TIR that consistently improves multi-turn TIR performance and tool invocation accuracy over strong RL baselines, highlighting the necessity of difficulty-aware credit assignment for effective policy optimization in tool-integrated LLM agents.

Yu-Can Guo, Xiaohan Wang, Miao Su et al. · 0 citations
Preprint Aug 2026

When Not to Imitate: Boundary-Aware Skill Memory for Reliable Tool-Use LLM Agents

BASM is proposed, which augments each skill with explicit boundary fields, which transforms each retrieved skill from an unconditional action template into state-conditioned guidance: the agent applies the skill when its conditions hold, suppresses inapplicable tool calls when they do not, and issues targeted repairs when execution fails.

Zi-Han Lin, Zhenyu Chen, Jiawen Wei et al. · 0 citations
Jul 2026

UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams

Inspired by the human brain, which balances plasticity and stability through complementary episodic storage and gradual consolidation, UniMem is proposed, a self-routing framework for autonomous memory management that consistently outperforms baselines while maintaining execution fidelity.

Si-Yu Xia, Chen-Heng Zhang, Yan-Ting Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.