Self-organizing state aggregation for architecture design of Q-learning

Kao-Shing Hwang,Hsin-Yi Lin,Yuan-Pao Hsu,Hung-Hsiu Yu

doi:10.1016/j.ins.2011.02.017

Abstract

This work describes a novel algorithm that integrates an adaptive resonance method (ARM), i.e. an ART-based algorithm with a self-organized design, and a Q-learning algorithm. By dynamically adjusting the size of sensitivity regions of each neuron and adaptively eliminating one of the redundant neurons, ARM can preserve resources, i.e. available neurons, to accommodate additional categories. As a dynamic programming-based reinforcement learning method, Q-learning involves use of the learned action-value function, Q, which directly approximates Q ∗, i.e. the optimal action-value function, which is independent of the policy followed. In the proposed algorithm, ARM functions as a cluster to classify input vectors from the outside world. Clustered results are then sent to the Q-learning design in order to learn how to implement the optimum actions to the outside world. Simulation results of the well-known control algorithm of balancing an inverted pendulum on a cart demonstrates the effectiveness of the proposed algorithm.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Self-organizing state aggregation for architecture design of Q-learning

Abstract

Talk to us

Similar Papers

More From: Information Sciences

Lead the way for us

Journal: Information Sciences	Publication Date: Mar 8, 2011
Citations: 42

Similar Papers

An ARM-Based Q-Learning Algorithm
Yuan-Pao Hsu ... Hsin-Yi Lin
-
Yuan-Pao Hsu, et. al.Yuan-Pao Hsu ... Hsin-Yi Lin
21 Aug 2007
21 Aug 2007

Reinforcement learning in the environment where optimal action value function is partly discontinuous
Shingo Shibusawa ... Takeshi Shibuya
-
Shingo Shibusawa, et. al.Shingo Shibusawa ... Takeshi Shibuya
01 Sep 2016
01 Sep 2016

Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar ... Rémi Munos
Machine Learning | VOL. 91
Mohammad Gheshlaghi Azar, et. al.Mohammad Gheshlaghi Azar ... Rémi Munos
14 May 2013
Machine Learning | VOL. 91

Adaptive attitude determination of bionic polarization integrated navigation system based on reinforcement learning strategy
Huiyi Bao ... Luyue Sun
Mathematical Foundations of Computing | VOL. 6
Huiyi Bao, et. al.Huiyi Bao ... Luyue Sun
01 Jan 2023
Mathematical Foundations of Computing | VOL. 6

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Self-organizing state aggregation for architecture design of Q-learning

Abstract

Talk to us

Similar Papers

More From: Information Sciences