A Rollout Algorithm for Multichain Markov Decision Processes with Average Cost

Tao Sun,Peter B Luh,Qianchuan Zhao

doi:10.1007/978-3-642-02894-6_15

Abstract

Many of simulation based learning algorithms have been developed to obtain near optimal policies for Markov decision processes (MDPs) with large state space. However, most of them are for unichain problems. In view that some applications involve multichain processes and it is NP-hard to determine whether a MDP is unichain or not, it is desirable to obtain an algorithm that is applicable to multichain problems as well. This paper presents a rollout algorithm for multichain MDPs with average cost. Preliminary analysis of the estimation error and parameter settings are provided based on the problem structures, i.e., mixing time of transition matrix. Ordinal optimization and Optimal Computing Budget Allocation are also suggested to improve the efficiency of the algorithm.

Full Text