Offline reinforcement learning in high-dimensional stochastic environments

Félicien Hêche,Oussama Barakat,Thibaut Desmettre,Tania Marx,Stephan Robert-Nicoud

doi:10.1007/s00521-023-09029-3

Abstract

Offline reinforcement learning (RL) has emerged as a promising paradigm for real-world applications since it aims to train policies directly from datasets of past interactions with the environment. The past few years, algorithms have been introduced to learn from high-dimensional observational states in offline settings. The general idea of these methods is to encode the environment into a latent space and train policies on top of this smaller representation. In this paper, we extend this general method to stochastic environments (i.e., where the reward function is stochastic) and consider a risk measure instead of the classical expected return. First, we show that, under some assumptions, it is equivalent to minimizing a risk measure in the latent space and in the natural space. Based on this result, we present Latent Offline Distributional Actor-Critic (LODAC), an algorithm which is able to train policies in high-dimensional stochastic and offline settings to minimize a given risk measure. Empirically, we show that using LODAC to minimize Conditional Value-at-Risk (CVaR) outperforms previous methods in terms of CVaR and return on stochastic environments.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Offline reinforcement learning in high-dimensional stochastic environments

Abstract

Talk to us

Similar Papers

More From: Neural Computing and Applications

Lead the way for us

Journal: Neural Computing and Applications	Publication Date: Oct 11, 2023
License type: CC BY 4.0

Similar Papers

Innovative transition matrix techniques for measuring extreme risk: an Australian and U.S. comparison
...
-
, et. al. ...
01 Jan 2010
01 Jan 2010

Reinforcement Learning in Latent Action Sequence Space
Heecheol Kim ... Kosuke Miyoshi
-
Heecheol Kim, et. al.Heecheol Kim ... Kosuke Miyoshi
24 Oct 2020
24 Oct 2020

Taming Tail Risk: Regularized Multiple β Worst-Case CVaR Portfolio
Kei Nakagawa ... Katsuya Ito
Symmetry | VOL. 13
Kei Nakagawa, et. al.Kei Nakagawa ... Katsuya Ito
21 May 2021
Symmetry | VOL. 13

Multi-Group Transfer Learning on Multiple Latent Spaces for Text Classification
Jianhan Pan ... Thuc Duy Le
IEEE Access | VOL. 8
Jianhan Pan, et. al.Jianhan Pan ... Thuc Duy Le
01 Jan 2020
IEEE Access | VOL. 8

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Offline reinforcement learning in high-dimensional stochastic environments

Abstract

Talk to us

Similar Papers

More From: Neural Computing and Applications