Fast Two-Stage Computation of an Index Policy for Multi-Armed Bandits with Setup Delays

José Niño-Mora

doi:10.3390/math9010052

Abstract

We consider the multi-armed bandit problem with penalties for switching that include setup delays and costs, extending the former results of the author for the special case with no switching delays. A priority index for projects with setup delays that characterizes, in part, optimal policies was introduced by Asawa and Teneketzis in 1996, yet without giving a means of computing it. We present a fast two-stage index computing method, which computes the continuation index (which applies when the project has been set up) in a first stage and certain extra quantities with cubic (arithmetic-operation) complexity in the number of project states and then computes the switching index (which applies when the project is not set up), in a second stage, with quadratic complexity. The approach is based on new methodological advances on restless bandit indexation, which are introduced and deployed herein, being motivated by the limitations of previous results, exploiting the fact that the aforementioned index is the Whittle index of the project in its restless reformulation. A numerical study demonstrates substantial runtime speed-ups of the new two-stage index algorithm versus a general one-stage Whittle index algorithm. The study further gives evidence that, in a multi-project setting, the index policy is consistently nearly optimal.

Highlights

In a much-studied version of the multi-armed bandit problem (MABP), a decision-maker selects one project to engage from a finite set of dynamic and stochastic projects at each of an infinite sequence of discrete-time periods
When a Markovian non-restless bandit with switching delays is reformulated as a semi-Markov restless bandit without them, it is found that the resultant model need not satisfy the partial conservation laws (PCLs)-indexability conditions that were the cornerstone to the analyses presented in Niño-Mora [27] for the pure-switching-costs case
Concerning the second goal, on general restless bandit methodology, we introduce, for finite-state restless bandits, significantly simpler and less stringent sufficient conditions for indexability than the former PCL-based conditions, under which it is assured that the adaptive-greedy algorithm computes the MPI

Summary

Introduction

In a much-studied version of the multi-armed bandit problem (MABP), a decision-maker selects one project to engage from a finite set of dynamic and stochastic projects at each of an infinite sequence of discrete-time periods. Each project is modeled as a classic (non-restless) bandit, so the engaged (active) project gives rewards and its state changes in a Markovian fashion, while rested (passive) projects neither produce rewards nor change state. The goal is to find a policy that selects one project to be engaged at each time, for maximizing the expected total geometrically discounted reward. The MABP is widely applicable, being regarded as a modeling paradigm of the exploration versus exploitation trade-off, and it has generated a vast literature (see the monograph [1] and the cited references there). The curse of dimensionality hinders direct numerical solution of its dynamic programming (DP)

Objectives

Results

Conclusion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Mathematics	Publication Date: Dec 29, 2020
Citations: 2	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Fast Two-Stage Computation of an Index Policy for Multi-Armed Bandits with Setup Delays

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Mathematics

Lead the way for us

Similar Papers

Asymptotically optimal index policies for an abandonment queue with convex holding cost
M Larrañaga ... I M Verloop
Queueing Systems | VOL. 81
M Larrañaga, et. al.M Larrañaga ... I M Verloop
14 May 2015
Queueing Systems | VOL. 81

Downlink Scheduling Over Markovian Fading Channels
Wenzhuo Ouyang ... Atilla Eryilmaz
IEEE/ACM Transactions on Networking | VOL. 24
Wenzhuo Ouyang, et. al.Wenzhuo Ouyang ... Atilla Eryilmaz
01 Jun 2016
IEEE/ACM Transactions on Networking | VOL. 24

Asymptotically optimal downlink scheduling over Markovian fading channels
Wenzhuo Ouyang ... Atilla Eryilmaz
-
Wenzhuo Ouyang, et. al.Wenzhuo Ouyang ... Atilla Eryilmaz
01 Mar 2012
01 Mar 2012

Channel probing for opportunistic access with multi-channel sensing
Keqin Liu ... Qing Zhao
-
Keqin Liu, et. al.Keqin Liu ... Qing Zhao
01 Oct 2008
01 Oct 2008

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Fast Two-Stage Computation of an Index Policy for Multi-Armed Bandits with Setup Delays

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Mathematics