Similar papers
Understanding cyclist route choice: A hierarchical adversarial inverse reinforcement learning approach to behavioral profiling
Reinforcement learning optimization of a ⁶Li magneto-optical trap: DDPG versus SAC
Key Developments in Reinforcement Learning in Robotics and Practical Implications: A Comprehensive Survey
This narrative review examines the current state of knowledge regarding Reinforcement Learning in Robotics within the broader context of Artificial Intelligence. We survey the theoretical foundations, methodological approaches, and key findings that have shaped the field, identifying major themes and tracing the evolution of ideas over time. The review synthesizes evidence from multiple research traditions and highlights both established conclusions and areas of ongoing debate. Particular attention is given to recent advances that have opened new avenues for investigation and to the practical implications of theoretical developments. We conclude with a discussion of the most promising directions for future research, emphasizing the importance of interdisciplinary collaboration and methodological innovation.
AgentCreditBench: A Conformance-Test Suite for Turn-Level Credit Estimators
AgentCreditBench is a CPU-first conformance-test suite for turn-level credit assignment in agentic reinforcement learning. It compares GRPO, RLOO, GAE, GiGPO, Monte Carlo, and custom estimator outputs with exact policy advantages on tiny finite-horizon Markov decision processes, and separately evaluates the induced policy-gradient signal.
Learning by Watching: A Narrative Review of Imitation Learning from ALVINN to Generative Adversarial Imitation
Imitation learning---the learning of behavior from demonstrations instead of rewards---moved from Pomerleau's ALVINN driving network and Schaal's humanoid route through Ng and Russell's inverse reinforcement learning, Abbeel and Ng's apprenticeship learning, and Ziebart's maximum entropy to the robot learning from demonstration surveys, Ross's DAgger, Ho and Ermon's generative adversarial imitation, Finn's guided cost learning, and the algorithmic perspective's syntheses. This article presents a narrative review of that arc's canonical line: Pomerleau's 1989 ALVINN, Schaal's 1999 humanoid question, Ng and Russell's 2000 inverse RL, Abbeel and Ng's 2004 apprenticeship learning, Billard and colleagues's 2008 handbook chapter, Ziebart and colleagues's 2008 maximum entropy, Argall and colleagues's 2009 survey, Ross, Gordon, and Bagnell's 2011 DAgger, Ho and Ermon's 2016 GAIL, Finn and colleagues's 2016 guided cost learning, Hussein and colleagues's 2017 survey, and Osa and colleagues's 2018 algorithmic perspective. The review is organized around three themes: the foundations, in which the driving network's demonstrations, the humanoid's question, and the inverse reward's recovery defined the field's two programs; the demonstration's surveys, in which the robot programming's handbook and the LfD's survey systematized the practice; and the deep era, in which the DAgger's covariate correction, the adversarial's discrimination, and the algorithmic perspective's synthesis unified the field. It is concluded that imitation learning is the reward's workaround---and that its arc is the demonstrator's knowledge's transfer from the human's steering to the policy's distributions.
When to Parallelize Stochastic Exploration of Rare Rewards in Reinforcement Learning
International audience
Related blog posts
Looking beyond natural sequences
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
Generating scenarios for extreme events, without extreme data
A new algorithm learns to anticipate the unprecedented scenarios that critical infrastructure and global supply chains are least prepared for.
When AI art has no author: Study finds generated images often can’t be traced to training data
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
What We Learned by Reproducing 2,200 papers from ICML
We’re on a journey to advance and democratize artificial intelligence through open source and open science.