Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Jun 2025

BOW: Training Language Models to Reason Over Plausible Next Words

BOW is introduced, an RL framework that instead trains models to produce self-contained, neutral, and comprehensive descriptions of the plausible next-word space, and human evaluation shows that BOW-Reg produces broader next-word reasoning trajectories, while direct next-word-prediction evaluation shows that these trajectories remain predictive.

Ming Shen, Zhikun Xu, Xiao Ye et al. · 2 citations

Skill Reuse as Compression in Agentic RL

This work proves a PAC-Bayes bound guaranteeing that a dictionary extracted from successful trajectories has bounded expected description length on future successful behavior, and introduces ReuseRL, which grounds agentic RL in the Minimum Description Length (MDL) principle.

Zhikun Xu, Yu Feng, Jacob Dineen et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.