Approximate regret based elicitation in Markov decision process

Pegah Alizadeh,Yann Chevaleyre,Jean-Daniel Zucker

doi:10.1109/rivf.2015.7049873

Abstract

Consider a decision support system (DSS) designed to find optimal strategies in stochastic environments, on behalf of a user. To perform this computation, the DSS will need a precise model of the environment. Of course, when the environment can be modeled as a Markov decision process (MDP) with numerical rewards (or numerical penalties), the DSS can compute the optimal strategy in polynomial time. But in many real-world cases, rewards are unknown. To compensate this missing information, the DSS may query the user for its preferences among some alternative policies. Based on the user's answers, the DSS can step-by-step compute the user's preferred policy. In this work, we describe a computational method based on minimax regret to find optimal policy when rewards are unknown. Then we present types of queries on feasible set of rewards by using preference elicitation approaches. When user answers these queries based on her preferences, we will have more information about rewards which will result in more desirable policies.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Approximate regret based elicitation in Markov decision process

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Application of Markov decision processes to search problems
Leo B Hartman ... Kees M Van Hee
Decision Support Systems | VOL. 14
Leo B Hartman, et. al.Leo B Hartman ... Kees M Van Hee
01 Jul 1995
Decision Support Systems | VOL. 14

The integrated control of production-inventory systems

-

06 Dec 2005
06 Dec 2005

Equivalence notions and model minimization in Markov decision processes
Robert Givan ... Matthew Greig
Artificial Intelligence | VOL. 147
Robert Givan, et. al.Robert Givan ... Matthew Greig
12 Feb 2003
Artificial Intelligence | VOL. 147

RL Based Decision Support System for u-Healthcare Environment
Devinder Thapa ... Gi-Nam Wang
-
Devinder Thapa, et. al.Devinder Thapa ... Gi-Nam Wang
01 Jan 2008
01 Jan 2008

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Approximate regret based elicitation in Markov decision process

Abstract

Talk to us

Similar Papers