Skip to content
Open access

Parallel Heuristic Search as Inference for Actor-Critic Reinforcement Learning Models (Extended Abstract)

Aug 2026 · Proceedings of the International Symposium on Combinatorial Search · 0 citations

Abstract

Actor-critic models are a class of model-free deep reinforcement learning (RL) algorithms that have demonstrated effectiveness across various robot learning tasks. While considerable research has focused on improving training stability and data sampling efficiency, most deployment strategies have remained relatively simplistic, typically relying on direct actor policy rollouts. In contrast, we propose PACHS (Parallel Actor-Critic Heuristic Search), an efficient parallel best-first search algorithm for inference that leverages both components of the actor-critic architecture: the actor network generates actions, while the critic network provides cost-to-go estimates to guide the search. Two levels of parallelism are employed within the search---actions and cost-to-go estimates are generated in batches by the actor and critic networks respectively, and graph expansion is distributed across multiple threads. We demonstrate the effectiveness of our approach in robotic manipulation tasks, including collision-free motion planning and contact-rich interactions such as non-prehensile pushing. Visit https://p-achs.github.io for demonstrations and examples.

Read PDF

Similar papers

Preprint Aug 2026

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition

A unified framework that combines centralized training with decentralized execution (CTDE) and a Hybrid Reward Architecture (HRA) is introduced that enables multiple actors to share a centralized multi-head critic and substantially improves both sample efficiency and policy performance.

Changhao Li, Yifang Zhang, Heng Zhang et al. · 0 citations
Preprint Sep 2026

GR2PO: Group Relative Return Policy Optimization for Continuous Robot Control

Actor-critic architecture has been widely used in continuous robot control. However, they rely on learning a value network, introducing additional computational overhead during training. Moreover, policy learning may also be affected by the approximation error of value estimation. Critic-free group relative policy opti...

Pengqin Wang, Qi-Ming Zhang, Shao-Jie Shen et al. · 0 citations
Preprint Aug 2026

Decoupling Policy Extraction for Offline Reinforcement Learning

Offline RL methods commonly jointly train the actor and critic, where the critic is used to guide the actor toward higher-value actions. This coupled learning process is well motivated in online RL, where an improved actor collects new data that can further update the actor and the critic. However, training data remain...

Xu-Yao Lin, Yixiang Shan, Jin-Ru Duan et al. · 0 citations
#machine learning Preprint Aug 2026

Locally-Guided Actor-Critic: Training a Goal-conditioned Actor with a Subgoal-aware Critic

An alternative reward shaping method (RS) is proposed that removes deceptive rewards at the expense of theoretical guarantees of PBRS, and another method named Locally-Guided Actor Critic (LG-AC) that rewards the agent for reaching intermediate goals achieves the best overall performance across tasks.

Olivier Serris, Stéphane Doncieux, Olivier Sigaud · 0 citations
Book Open access Aug 2026

Large Language Model (LLM) as an Excellent Reinforcement Learning Researcher in both Single-Agent and Multi-Agent Scenarios

A Self-Evolutional single-agent/multi-agent Reinforcement Learning (SE-RL) framework that utilizes a Large Language Model (LLM) to design various RL algorithm modules, such as agent model design, reward function, profiling, communication, and state imagination, by leveraging the LLM generating module output or code.

Vincent Fu, Xin-Xin Xu, Weichen Xu et al. · 0 citations
Preprint Sep 2026

CAST: Alternating State-Value Targets and Expanded Policy Gradients for Model-Based Reinforcement Learning

Model-based reinforcement learning (MBRL) is a family of RL methods that learn a model of the environment and use it for action selection, making it well suited to robotics due to its sample efficiency. Combining learned models with online planning can further improve action selection, as the planner can exploit the mo...

Pietro Noah Crestaz, Mohamed Yassine Kabouri, Nicolas Mansard et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.