Skip to content

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Oct 2026

Programmatic Search Agents: Extending Agentic Search Beyond Query Reformulation

Search agents adapt their queries, yet fixed search interfaces leave candidate processing and evidence presentation outside the agent's direct control. Our trajectory analysis shows that supporting passages can be retrieved yet never delivered to the agent; a same-page oracle intervention shows that changing the return...

Jia-Ming Qian, Hui-Yan Yang, Man-Di Liu et al. · 0 citations
#machine learning Preprint Sep 2026

RMB: Reward Model Boosting Mitigates Reward Hacking

Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning large language models (LLMs) with human preference. However, it often suffers from the reward hacking issue, where policy optimization improves the proxy reward model while actually degrading performance with respect to the true huma...

Jia-Bin Fan, De-Zhi Ye, Yong-Chang Hao et al. · 0 citations
Book Open access Aug 2026

UniRank: A Unified Framework for Efficient Multi-Objective LLM Ranking in Industrial Search

Multi-objective ranking serves as the backbone of industrial information retrieval, requiring a holistic assessment of documents across dimensions such as Relevance, Authority, and Recency. The prevailing industry paradigm relies on ensembles of specialized BERT-based models, which are costly to maintain and fundamenta...

De-Zhi Ye, Junwei Hu, Xiaoyang Chen et al. · 0 citations
#natural language process... Preprint Aug 2026

Beyond Polarization: The Generative Constraint of Chain-of-Thought in Pointwise Reranking

These findings suggest that, within the pointwise scoring paradigm, routing continuous relevance semantics through discrete text constrains ranking signal resolution reveals a bottleneck that is stable and difficult to overcome under current standard methods, rather than an easily resolvable training bias.

Xiaoyang Chen, Jie Liu, Haijin Liang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.