Reward Shaping with Dynamic Trajectory Aggregation

Takato Okudo,Seiji Yamada

doi:10.1109/ijcnn52387.2021.9533401

Abstract

Reinforcement learning, which acquires a policy maximizing long-term rewards, has been actively studied. Unfortunately, this learning type is too slow and difficult to use in practical situations because the state-action space becomes huge in real environments. The essential factor for learning efficiency is rewards. Potential-based reward shaping is a basic method for enriching rewards. This method is required to define a specific real-value function called a “potential function” for every domain. It is often difficult to represent the potential function directly. SARSA-RS learns the potential function and acquires it. However, SARSA-RS can only be applied to the simple environment. The bottleneck of this method is the aggregation of states to make abstract states since it is almost impossible for designers to build an aggregation function for all states. We propose a trajectory aggregation that uses subgoal series. This method dynamically aggregates states in an episode during trial and error with only the subgoal series and subgoal identification function. It makes designer effort minimal and the application to environments with high-dimensional observations possible. We obtained subgoal series from participants for experiments. We conducted the experiments in three domains, four-rooms(discrete states and discrete actions), pinball(continuous and discrete), and picking(both continuous). We compared our method with a baseline reinforcement learning algorithm and other subgoal-based methods, including random subgoal and naive subgoal-based reward shaping. As a result, our reward shaping outperformed all other methods in learning efficiency.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Reward Shaping with Dynamic Trajectory Aggregation

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Subgoal-Based Reward Shaping to Improve Efficiency in Reinforcement Learning
Takato Okudo ... Seiji Yamada
IEEE Access | VOL. 9
Takato Okudo, et. al.Takato Okudo ... Seiji Yamada
01 Jan 2020
IEEE Access | VOL. 9

Learning Potential in Subgoal-Based Reward Shaping
Takato Okudo ... Seiji Yamada
IEEE Access | VOL. 11
Takato Okudo, et. al.Takato Okudo ... Seiji Yamada
01 Jan 2023
IEEE Access | VOL. 11

Reward shaping using directed graph convolution neural networks for reinforcement learning and games
Jianghui Sang ... Hengfu Yin
Frontiers in Physics | VOL. 11
Jianghui Sang, et. al.Jianghui Sang ... Hengfu Yin
09 Nov 2023
Frontiers in Physics | VOL. 11

Policy invariant explicit shaping: an efficient alternative to reward shaping
Paniz Behboudian ... Yash Satsangi
Neural Computing and Applications | VOL. 34
Paniz Behboudian, et. al.Paniz Behboudian ... Yash Satsangi
28 Sep 2021
Neural Computing and Applications | VOL. 34

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Reward Shaping with Dynamic Trajectory Aggregation

Abstract

Talk to us

Similar Papers