A three-paradigm taxonomy (feature-based, auxiliary-based, and policy-based) based on the functional role of LLMs within the RL pipeline is proposed, which provides superior scalability and stability, though often at the expense of representational depth.
Ghusoon Hadi al-Aldaffaie, Alireza Taheri, Amirfarhad Farhadi et al.· Discover Artificial Intellig...· 0 citations
: Deep reinforcement learning (DRL) has made significant progress in recent years. However, it still faces multiple bottlenecks in real-world applications. Agents suffer from low sample efficiency, and designing effective reward functions remains very difficult. Furthermore, traditional DRL lacks general cognitive abil...
Sheng-Hao Yuan· Proceedings of the 4th Inter...· 0 citations
Two-tiered organization provides a compact survey lens for autonomous control (AC): an upper tier supports learning, planning, and high-level (HL) mission reasoning, while a lower tier executes feedback control, safety filtering, and platform-specific constraints. This compact survey reviews hierarchical architectures...
Junseo Min, Hoyeong Lee, Hyojun Ahn et al.· International Conference on...· 0 citations
As cyber threats continue to evolve, there is a need for Autonomous Cyber Defense (ACD) strategies capable of fast and context-aware responses. Reinforcement learning (RL) has shown promise in automating cyber defense by exploring and learning effective countermeasures. However, RL often struggles with sparse reward si...
Md. Shamim Towhid, Shahrear Iqbal, E. C. Pinto et al.· IEEE Transactions on Network...· 0 citations
A structured review of three key roles that RL plays in empowering OR, serving as an end-to-end solution method or as a component integrated within heuristic and exact OR methods for combinatorial optimization problems, and facilitating extended reality analysis through integration with digital twin systems is presente...
Ya-Han Lu, Dong-Yang Xia, Nurşen Aydın et al.· 0 citations
DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization) integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training.
Young Kyu Yu, Sanghwan Jang, Hwanjo Yu· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.