Skip to content
Open access

Large Language Models and Reinforcement Learning: A Taxonomy of Integration Paradigms, Challenges, and Future Directions

Aug 2026 · Computers and artificial intelligence · 0 citations · 26 references

Abstract

This paper examines the integration of large language models (LLMs) and reinforcement learning (RL) in recommender systems, focusing on their theoretical foundations and structural challenges. It highlights the transition from static prediction to sequential decision-making, emphasizing RL’s strengths in long-term reward optimization and interaction modeling, and LLMs’ advantages in semantic understanding and reasoning. Their complementary limitations—RL’s weak semantic representation and LLMs’ lack of long-term optimization—justify their integration. Existing research is classified into “LLM-enhanced RL” and “RL-shaped LLM,” with roles including representation enhancement, reward modeling, policy generation, and environment simulation, under varying coupling levels. The paper proposes a unified three-dimensional framework based on information sources, optimization time scale, and coupling strength, showing that performance differences arise from structural positioning rather than model scale. Key challenges include balancing expressiveness and efficiency, long-term optimization and training stability, and generalization versus specialization. The paper also identifies limitations in evaluation protocols and experimental design, calling for standardized frameworks for long-term value assessment. Overall, integrating LLMs and RL is crucial for advancing recommender systems toward intelligent decision-making agents, with future work focusing on stable coupling and unified evaluation.

Read PDF