Constrained Reinforcement Learning in Hard Exploration Problems

Pathmanathan Pankayaraj,Pradeep Varakantham

doi:10.1609/aaai.v37i12.26757

Abstract

One approach to guaranteeing safety in Reinforcement Learning is through cost constraints that are dependent on the policy. Recent works in constrained RL have developed methods that ensure constraints are enforced even at learning time while maximizing the overall value of the policy. Unfortunately, as demonstrated in our experimental results, such approaches do not perform well on complex multi-level tasks, with longer episode lengths or sparse rewards. To that end, we propose a scalable hierarchical approach for constrained RL problems that employs backward cost value functions in the context of task hierarchy and a novel intrinsic reward function in lower levels of the hierarchy to enable cost constraint enforcement. One of our key contributions is in proving that backward value functions are theoretically viable even when there are multiple levels of decision making. We also show that our new approach, referred to as Hierarchically Limited consTraint Enforcement (HiLiTE) significantly improves on state of the art Constrained RL approaches for many benchmark problems from literature. We further demonstrate that this performance (on value and constraint enforcement) clearly outperforms existing best approaches for constrained RL and hierarchical RL.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Constrained Reinforcement Learning in Hard Exploration Problems

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Similar Papers

Gradient-Adaptive Pareto Optimization for Constrained Reinforcement Learning
Zixian Zhou ... Jia He
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 37
Zixian Zhou, et. al.Zixian Zhou ... Jia He
26 Jun 2023
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 37

Author response: On the normative advantages of dopamine and striatal opponency for learning and choice
Alana Jaskir ... Michael J Frank
-
Alana Jaskir, et. al.Alana Jaskir ... Michael J Frank
14 Feb 2023
14 Feb 2023

A Modern Perspective on Safe Automated Driving for Different Traffic Dynamics Using Constrained Reinforcement Learning
Danial Kamran ... Matthijs T.J Spaan
-
Danial Kamran, et. al.Danial Kamran ... Matthijs T.J Spaan
08 Oct 2022
08 Oct 2022

Online Model Learning Algorithms for Actor-Critic Control

-

04 Mar 2015
04 Mar 2015

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Constrained Reinforcement Learning in Hard Exploration Problems

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence