MM-KTD: Multiple Model Kalman Temporal Differences for Reinforcement Learning

Parvin Malekzadeh,Arash Mohammadi,Mohammad Salimibeni,Akbar Assa,Konstantinos N Plataniotis

doi:10.1109/access.2020.3007951

Parvin Malekzadeh, Arash Mohammadi + Show 3 more

Open Access

https://doi.org/10.1109/access.2020.3007951

Copy DOI

Journal: IEEE Access	Publication Date: Jan 1, 2020
Citations: 38	License type: CC BY 4.0

Affiliation: Concordia University, University of Toronto

Abstract

Background : There has been an increasing surge of interest on development of advanced Reinforcement Learning (RL) systems as intelligent approaches to learn optimal control policies directly from smart agents’ interactions with the environment. Objectives : In a model-free RL method with continuous state-space, typically, the value function of the states needs to be approximated. In this regard, Deep Neural Networks (DNNs) provide an attractive modeling mechanism to approximate the value function using sample transitions. DNN-based solutions, however, suffer from high sensitivity to parameter selection, are prone to overfitting, and are not very sample efficient. A Kalman-based methodology, on the other hand, could be used as an efficient alternative. Such an approach, however, commonly requires a-priori information about the system (such as noise statistics) to perform efficiently. The main objective of this paper is to address this issue. Methods : As a remedy to the aforementioned problems, this paper proposes an innovative Multiple Model Kalman Temporal Difference (MM-KTD) framework, which adapts the parameters of the filter using the observed states and rewards. Moreover, an active learning method is proposed to enhance the sampling efficiency of the system. More specifically, the estimated uncertainty of the value functions are exploited to form the behaviour policy leading to more visits to less certain values, therefore, improving the overall learning sample efficiency. As a result, the proposed MM-KTD framework can learn the optimal policy with significantly reduced number of samples as compared to its DNN-based counterparts. Results : To evaluate performance of the proposed MM-KTD framework, we have performed a comprehensive set of experiments based on three RL benchmarks, namely, Inverted Pendulum; Mountain Car, and; Lunar Lander. Experimental results show superiority of the proposed MM-KTD framework in comparison to its state-of-the-art counterparts.

Highlights

Inspired by exceptional learning capabilities of human beings, Reinforcement Learning (RL) systems have emerged aiming to form optimal control policies merely by relying on the knowledge about the past interactions of an agent with its environment
EXPERIMENTAL RESULTS we evaluate performance of the proposed Multiple Model Kalman Temporal Difference (MM-Kalman Temporal difference (KTD)) framework
It is worth mentioning that one benefit of the proposed MM-KTD framework is its superior ability to deal with scenarios where enough information about the underlying parameters is not fully available

Summary

Introduction

Inspired by exceptional learning capabilities of human beings, Reinforcement Learning (RL) systems have emerged aiming to form optimal control policies merely by relying on the knowledge about the past interactions of an agent with its environment. Such a learning approach is beneficial, as unlike supervised learning methods, an RL system. Methods: As a remedy to the aforementioned problems, this paper proposes an innovative Multiple Model Kalman Temporal Difference (MM-KTD) framework, which adapts the parameters of the filter using the observed states and rewards. The proposed MM-KTD framework can learn the optimal policy with significantly reduced number of samples as compared to its DNN-based counterparts. Experimental results show superiority of the proposed MM-KTD framework in comparison to its state-ofthe-art counterparts

Objectives

Results

Conclusion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

MM-KTD: Multiple Model Kalman Temporal Differences for Reinforcement Learning

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: IEEE Access

Lead the way for us

Similar Papers

Distributed Hybrid Kalman Temporal Differences for Reinforcement Learning
Mohammad Salimibeni ... Arash Mohammadi
-
Mohammad Salimibeni, et. al.Mohammad Salimibeni ... Arash Mohammadi
01 Nov 2020
01 Nov 2020

Artificial Intelligence and the Common Sense of Animals.
Murray Shanahan ... Matthew Crosby
Trends in Cognitive Sciences | VOL. 24
Murray Shanahan, et. al.Murray Shanahan ... Matthew Crosby
08 Oct 2020
Trends in Cognitive Sciences | VOL. 24

Efficient actor-critic algorithm with dual piecewise model learning
Shan Zhong ... Shengrong Gong
-
Shan Zhong, et. al.Shan Zhong ... Shengrong Gong
01 Nov 2017
01 Nov 2017

Interpretable policies for reinforcement learning by genetic programming
Daniel Hein ... Thomas A Runkler
Engineering Applications of Artificial Intelligence | VOL. 76
Daniel Hein, et. al.Daniel Hein ... Thomas A Runkler
22 Sep 2018
Engineering Applications of Artificial Intelligence | VOL. 76

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

MM-KTD: Multiple Model Kalman Temporal Differences for Reinforcement Learning

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: IEEE Access