Protecting Reward Function of Reinforcement Learning via Minimal and Non-catastrophic Adversarial Trajectory

Tong Chen,Yunzhe Tian,Gang Li,Wenjia Niu,Yike Li,Endong Tong,Jiqiang Liu,Yingxiao Xiang,Qi Alfred Chen

doi:10.1109/srds53918.2021.00037

Abstract

Reward functions are critical hyperparameters with commercial values for individual or distributed reinforcement learning (RL), as slightly different reward functions result in significantly different performance. However, existing inverse reinforcement learning (IRL) methods can be utilized to approximate reward functions just based on collected expert trajectories through observing. Thus, in the real RL process, how to generate a polluted trajectory and perform an adversarial attack on IRL for protecting reward functions has become the key issue. Meanwhile, considering the actual RL cost, generated adversarial trajectories should be minimal and non-catastrophic for ensuring normal RL performance. In this work, we propose a novel approach to craft adversarial trajectories disguised as expert ones, for decreasing the IRL performance and realize the anti-IRL ability. Firstly, we design a reward clustering-based metric to integrate both advantages of fine- and coarse-grained IRL assessment, including expected value difference (EVD) and mean reward loss (MRL). Further, based on such metric, we explore an adversarial attack based on agglomerative nesting algorithm (AGNES) clustering and determine targeted states as starting states for reward perturbation. Then we employ the intrinsic fear model to predict the probability of imminent catastrophe, supporting to generate non-catastrophic adversarial trajectories. Extensive experiments of 7 state-of-the-art IRL algorithms are implemented on the Object World benchmark, demonstrating the capability of our proposed approach in (a) decreasing the IRL performance and (b) having minimal and non-catastrophic adversarial trajectories.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Protecting Reward Function of Reinforcement Learning via Minimal and Non-catastrophic Adversarial Trajectory

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Sparse online maximum entropy inverse reinforcement learning via proximal optimization and truncated gradient
Li Song ... Xin Xu
Knowledge-Based Systems | VOL. 252
Li Song, et. al.Li Song ... Xin Xu
16 Jul 2022
Knowledge-Based Systems | VOL. 252

Inverse reinforcement learning using Dynamic Policy Programming
Eiji Uchibe ... Kenji Doya
-
Eiji Uchibe, et. al.Eiji Uchibe ... Kenji Doya
01 Oct 2014
01 Oct 2014

Inverse-Inverse Reinforcement Learning. How to Hide Strategy from an Adversarial Inverse Reinforcement Learner
Kunal Pattanayak ... Christopher Berry
-
Kunal Pattanayak, et. al.Kunal Pattanayak ... Christopher Berry
06 Dec 2022
06 Dec 2022

Reinforcement Learning for Clinical Applications.
Kia Khezeli ... Benjamin Shickel
Clinical journal of the American Society of Nephrology : CJASN | VOL. 18
Kia Khezeli, et. al.Kia Khezeli ... Benjamin Shickel
08 Feb 2023
Clinical journal of the American Society of Nephrology : CJASN | VOL. 18

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Protecting Reward Function of Reinforcement Learning via Minimal and Non-catastrophic Adversarial Trajectory

Abstract

Talk to us

Similar Papers