Omega-Regular Decision Processes

Ernst Moritz Hahn,Fabio Somenzi,Ashutosh Trivedi,Sven Schewe,Dominik Wojtczak,Mateo Perez

doi:10.1609/aaai.v38i19.30105

Abstract

Regular decision processes (RDPs) are a subclass of non-Markovian decision processes where the transition and reward functions are guarded by some regular property of the past (a lookback). While RDPs enable intuitive and succinct representation of non-Markovian decision processes, their expressive power coincides with finite-state Markov decision processes (MDPs). We introduce omega-regular decision processes (ODPs) where the non-Markovian aspect of the transition and reward functions are extended to an omega-regular lookahead over the system evolution. Semantically, these lookaheads can be considered as promises made by the decision maker or the learning agent about her future behavior. In particular, we assume that, if the promised lookaheads are not met, then the payoff to the decision maker is falsum (least desirable payoff), overriding any rewards collected by the decision maker. We enable optimization and learning for ODPs under the discounted-reward objective by reducing them to lexicographic optimization and learning over finite MDPs. We present experimental results demonstrating the effectiveness of the proposed reduction.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Omega-Regular Decision Processes

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Similar Papers

Probably Approximately Correct (PAC) exploration in reinforcement learning

-

01 Jan 2007
01 Jan 2007

Feature Reinforcement Learning: Part I. Unstructured MDPs
Marcus Hutter
Journal of Artificial General Intelligence | VOL. 1
Marcus HutterMarcus Hutter
01 Jan 2009
Journal of Artificial General Intelligence | VOL. 1

Feature Markov Decision Processes
Marcus Hutter
-
Marcus HutterMarcus Hutter
01 Jan 2009
01 Jan 2009

Process control using finite Markov chains with iterative clustering
Enso Ikonen ... István Selek
Computers & Chemical Engineering | VOL. 93
Enso Ikonen, et. al.Enso Ikonen ... István Selek
01 Jul 2016
Computers & Chemical Engineering | VOL. 93

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Omega-Regular Decision Processes

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence