Budget-Constrained Multi-Armed Bandits With Multiple Plays

Datong Zhou,Claire Tomlin

doi:10.1609/aaai.v32i1.11629

Abstract

We study the multi-armed bandit problem with multiple plays and a budget constraint for both the stochastic and the adversarial setting. At each round, exactly K out of N possible arms have to be played (with 1 ≤ K <= N). In addition to observing the individual rewards for each arm played, the player also learns a vector of costs which has to be covered with an a-priori defined budget B. The game ends when the sum of current costs associated with the played arms exceeds the remaining budget. Firstly, we analyze this setting for the stochastic case, for which we assume each arm to have an underlying cost and reward distribution with support [cmin, 1] and [0, 1], respectively. We derive an Upper Confidence Bound (UCB) algorithm which achieves O(NK4 log B) regret. Secondly, for the adversarial case in which the entire sequence of rewards and costs is fixed in advance, we derive an upper bound on the regret of order O(√NB log(N/K)) utilizing an extension of the well-known Exp3 algorithm. We also provide upper bounds that hold with high probability and a lower bound of order Ω((1 – K/N) √NB/K).

Full Text

Published version (

Free)

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Budget-Constrained Multi-Armed Bandits With Multiple Plays

Abstract

Talk to us

Similar Papers

More From: Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence	Publication Date: Apr 29, 2018
Citations: 14

Similar Papers

On adaptive estimation for dynamic Bernoulli bandits
Xue Lu ... Niall Adams
Foundations of data science (Springfield, Mo.) | VOL. 1
Xue Lu, et. al.Xue Lu ... Niall Adams
01 Jan 2019
Foundations of data science (Springfield, Mo.) | VOL. 1

Some Variations of Upper Confidence Bound for General Game Playing
Iván Francisco-Valencia ... José Raymundo Marcial-Romero
-
Iván Francisco-Valencia, et. al.Iván Francisco-Valencia ... José Raymundo Marcial-Romero
01 Jan 2019
01 Jan 2019

Upper Confidence Bound (UCB) Algorithms for Adaptive Operator Selection in MOEA/D
Richard A Gonçalves ... Carolina P Almeida
-
Richard A Gonçalves, et. al.Richard A Gonçalves ... Carolina P Almeida
01 Jan 2015
01 Jan 2015

Dynamic multi-arm bandit game based multi-agents spectrum sharing strategy design
Jingyang Lu ... Genshe Chen
-
Jingyang Lu, et. al.Jingyang Lu ... Genshe Chen
01 Sep 2017
01 Sep 2017

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Budget-Constrained Multi-Armed Bandits With Multiple Plays

Abstract

Talk to us

Similar Papers

More From: Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence