Fitted policy search

Martino Migliavacca,Andrea Bonarini,Alessio Pecorino,Marcello Restelli,Matteo Pirotta

doi:10.1109/adprl.2011.5967368

Abstract

In this paper we address the combination of batch reinforcement-learning (BRL) techniques with direct policy search (DPS) algorithms in the context of robot learning. Batch value-based algorithms (such as fitted Q-iteration) have been proved to outperform online ones in many complex applications, but they share the same difficulties in solving problems with continuous action spaces, such as robotic ones. In these cases, actor-critic and DPS methods are preferable, since the optimization process is limited to a family of parameterized (usually smooth) policies. On the other hand, these methods (e.g., policy gradient and evolutionary methods) are generally very expensive, since finding the optimal parameterization may require to evaluate the performance of several policies, which in many real robotic applications is unfeasible or even dangerous. To overcome such problems, we exploit the fitted policy search (FPS) approach, in which the expected return of any policy considered during the optimization process is evaluated offline (without resorting to the robot) by reusing the data collected in the initial exploration phase. In this way, it is possible to take the advantages of both BRL and DPS algorithms, thus achieving an effective learning approach to solve robotic problems. A balancing task on a real two-wheeled robotic pendulum is used to analyze the properties and evaluate the effectiveness of the FPS approach.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Fitted policy search

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

A Markov chain Monte Carlo algorithm for Bayesian policy search
Vahid Tavakol Aghaei ... Sinan Yıldırım
Systems Science & Control Engineering | VOL. 6
Vahid Tavakol Aghaei, et. al.Vahid Tavakol Aghaei ... Sinan Yıldırım
01 Jan 2018
Systems Science & Control Engineering | VOL. 6

Policy search with rare significant events: Choosing the right partner to cooperate with.
Paul Ecoffet ... Nicolas Fontbonne
PLOS ONE | VOL. 17
Paul Ecoffet, et. al.Paul Ecoffet ... Nicolas Fontbonne
26 Apr 2022
PLOS ONE | VOL. 17

Model-based direct policy search
...
-
, et. al. ...
10 May 2010
10 May 2010

Reinforcement Learning of Potential Fields to achieve Limit-Cycle Walking
Denise S Feirstein ... Heike Vallery
IFAC PapersOnLine | VOL. 49
Denise S Feirstein, et. al.Denise S Feirstein ... Heike Vallery
01 Jan 2015
IFAC PapersOnLine | VOL. 49

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Fitted policy search

Abstract

Talk to us

Similar Papers