Multi-step Reinforcement Learning Research Articles

Unifying seemingly disparate algorithmic ideas to produce better performing algorithms has been a longstanding goal in reinforcement learning. As a primary example, TD(λ) elegantly unifies one-step TD prediction with Monte Carlo methods through the use of eligibility traces and the trace-decay parameter. Currently, there are a multitude of algorithms that can be used to perform TD control, including Sarsa, Q-learning, and Expected Sarsa. These methods are often studied in the one-step case, but they can be extended across multiple time steps to achieve better performance. Each of these algorithms is seemingly distinct, and no one dominates the others for all problems. In this paper, we study a new multi-step action-value algorithm called Q(σ) that unifies and generalizes these existing algorithms, while subsuming them as special cases. A new parameter, σ, is introduced to allow the degree of sampling performed by the algorithm at each step during its backup to be continuously varied, with Sarsa existing at one extreme (full sampling), and Expected Sarsa existing at the other (pure expectation). Q(σ) is generally applicable to both on- and off-policy learning, but in this work we focus on experiments in the on-policy case. Our results show that an intermediate value of σ, which results in a mixture of the existing algorithms, performs better than either extreme. The mixture can also be varied dynamically which can result in even greater performance.

Read full abstract

Despite their proven effectiveness, many Michigan learning classifier systems (LCSs) cannot perform multistep reinforcement learning in continuous spaces. To meet this technical challenge, some LCSs have been designed to learn fuzzy logic rules. They can be largely classified into strength-based and accuracy-based systems. The latter is gaining more research attention in the last decade. However, existing accuracy-based learning systems either address primarily single-step learning problems or require the action space to be discrete. In this paper, a new accuracy-based learning fuzzy classifier system (LFCS) is developed to explicitly handle continuous state input and continuous action output during multistep reinforcement learning. Several technical improvements have been achieved while developing the new learning algorithm. Particularly, we have successfully extended ${Q}$ -learning like credit assignment methods to continuous spaces. To enable direct learning of stochastic strategies for action selection, we have also proposed to use a new fuzzy logic system with stochastic action outputs. Moreover, fine-grained learning of fuzzy rules has been achieved effectively in our algorithm by using a natural gradient learning method. It is the first time that these techniques are utilized substantially in any accuracy-based LFCSs. Meanwhile, in comparison with several recently proposed learning algorithms, our algorithm is shown to perform highly competitively on four benchmark learning problems and a robotics problem. The practical usefulness of our algorithm is also demonstrated by improving the performance of a wireless body area network.

Read full abstract

Multi-step Reinforcement Learning Research Articles

Related Topics

Articles published on Multi-step Reinforcement Learning

Automatic Generation Control in a Distributed Power Grid Based on Multi-Step Reinforcement Learning

MRL-Seg: Overcoming Imbalance in Medical Image Segmentation With Multi-Step Reinforcement Learning.

Learning Robust Predictive Control: A Spatial-Temporal Game Theoretic Approach.

Multi-step reinforcement learning for model-free predictive energy management of an electrified off-highway vehicle

Fully Convolutional Network with Multi-Step Reinforcement Learning for Image Processing

A novel multi-step reinforcement learning method for solving reward hacking

Multi-Step Reinforcement Learning: A Unifying Algorithm

Accuracy-Based Learning Classifier Systems for Multistep Reinforcement Learning: A Fuzzy Logic Approach to Handling Continuous Inputs and Learning Continuous Actions

Semiconductor final test scheduling with Sarsa( λ, k) algorithm

A Multi-Step Reinforcement Learning Algorithm

A parallel fuzzy inference model with distributed prediction scheme for reinforcement learning

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Multi-step Reinforcement Learning Research Articles

Related Topics

Articles published on Multi-step Reinforcement Learning

Automatic Generation Control in a Distributed Power Grid Based on Multi-Step Reinforcement Learning

MRL-Seg: Overcoming Imbalance in Medical Image Segmentation With Multi-Step Reinforcement Learning.

Learning Robust Predictive Control: A Spatial-Temporal Game Theoretic Approach.

Multi-step reinforcement learning for model-free predictive energy management of an electrified off-highway vehicle

Fully Convolutional Network with Multi-Step Reinforcement Learning for Image Processing

A novel multi-step reinforcement learning method for solving reward hacking

Multi-Step Reinforcement Learning: A Unifying Algorithm

Accuracy-Based Learning Classifier Systems for Multistep Reinforcement Learning: A Fuzzy Logic Approach to Handling Continuous Inputs and Learning Continuous Actions

Semiconductor final test scheduling with Sarsa( λ, k) algorithm

A Multi-Step Reinforcement Learning Algorithm

A parallel fuzzy inference model with distributed prediction scheme for reinforcement learning