Optimality of greedy policy for a class of standard reward function of restless multi-armed bandit problem

K Wang,L Chen,Q Liu

doi:10.1049/iet-spr.2011.0185

Abstract

In this paper,we consider the restless bandit problem, which is one of the most well-studied generalizations of the celebrated stochastic multi-armed bandit problem in decision theory. However, it is known be PSPACE-Hard to approximate to any non-trivial factor. Thus the optimality is very difficult to obtain due to its high complexity. A natural method is to obtain the greedy policy considering its stability and simplicity. However, the greedy policy will result in the optimality loss for its intrinsic myopic behavior generally. In this paper, by analyzing one class of so-called standard reward function, we establish the closed-form condition about the discounted factor \beta such that the optimality of the greedy policy is guaranteed under the discounted expected reward criterion, especially, the condition \beta = 1 indicating the optimality of the greedy policy under the average accumulative reward criterion. Thus, the standard form of reward function can easily be used to judge the optimality of the greedy policy without any complicated calculation. Some examples in cognitive radio networks are presented to verify the effectiveness of the mathematical result in judging the optimality of the greedy policy.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Optimality of greedy policy for a class of standard reward function of restless multi-armed bandit problem

Abstract

Talk to us

Similar Papers

More From: IET Signal Processing

Lead the way for us

Journal: IET Signal Processing	Publication Date: Jan 1, 2012
Citations: 24

Similar Papers

Approximation algorithms for restless bandit problems
...
-
, et. al. ...
04 Jan 2009
04 Jan 2009

Approximation Algorithms for Restless Bandit Problems
Sudipto Guha ... Kamesh Munagala
-
Sudipto Guha, et. al.Sudipto Guha ... Kamesh Munagala
04 Jan 2009
04 Jan 2009

Approximation algorithms for restless bandit problems
Sudipto Guha ... Kamesh Munagala
Journal of the ACM | VOL. 58
Sudipto Guha, et. al.Sudipto Guha ... Kamesh Munagala
01 Dec 2010
Journal of the ACM | VOL. 58

Optimal Management of Rechargeable Biosensors in Temperature-Sensitive Environments
Yahya Osais ... Marc St-Hilaire
-
Yahya Osais, et. al.Yahya Osais ... Marc St-Hilaire
01 Sep 2010
01 Sep 2010

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Optimality of greedy policy for a class of standard reward function of restless multi-armed bandit problem

Abstract

Talk to us

Similar Papers

More From: IET Signal Processing