Послiдовний розподiл ресурсiв устохастичному середовищi: загальнийопис, аналiз та чисельнi експерименти

A S Dzhoha

doi:10.17721/1812-5409.2021/3.1

A S Dzhoha

Open Access

PDF Available

https://doi.org/10.17721/1812-5409.2021/3.1

Copy DOI

Export

Save

Cite

Abstract
Highlights/Summary
Full-Text PDF
Similar Papers

Abstract

Listen

In this paper, we consider policies for the sequential resource allocation under the multi-armed bandit problem in a stochastic environment. In this model, an agent sequentially selects an action from a given set and an environment reveals a reward in return. In the stochastic setting, each action is associated with a probability distribution with parameters that are not known in advance. The agent makes a decision based on the history of the chosen actions and obtained rewards. The objective is to maximize the total cumulative reward, which is equivalent to the loss minimization. We provide a brief overview of the sequential analysis and an appearance of the multi-armed bandit problem as a formulation in the scope of the sequential resource allocation theory. Multi-armed bandit classification is given with an analysis of the existing policies for the stochastic setting. Two different approaches are shown to tackle the multi-armed bandit problem. In the frequentist view, the confidence interval is used to express the exploration-exploitation trade-off. In the Bayesian approach, the parameter that needs to be estimated is treated as a random variable. Shown, how this model can be modelled with help of the Markov decision process. In the end, we provide numerical experiments in order to study the effectiveness of these policies.

Highlights

У данiй статтi наводиться стислий огляд послiдовного аналiзу, а також мiсце в ньому послiдовного розподiлу ресурсiв
Each action is associated with a probability distribution with parameters that are not known in advance
We provide a brief overview of the sequential analysis and an appearance of the multi-armed bandit problem as a formulation in the scope of the sequential resource allocation theory

Summary

Класифiкацiя моделей багаторукого бандита

Проблема багаторуких бандитiв це послiдовна взаємодiя мiж суб’єктом, що приймає рiшення, так званим агентом, та зовнiшнiм середовищем. Метою агента є послiдовний вибiр таких дiй iз заданої множини A, якi призводять до найбiльшої можливої сукупної винагороди за n крокiв, тобто n t=1. Ключовим моментом у данiй проблемi є те, що середовище не вiдоме для агента, тобто не вiдомий розподiл iмовiрностей винагород кожної дiї моделi. Однiєю з метрик вимiрювання ефективностi стратегiї агента є його втрати Rn за n крокiв, що є рiзницею мiж очiкуваною винагородою при виборi оптимальної дiї на кожному кроцi та винагородою при виборi дiй вiдповiдно до заданої стратегiї. За характером взаємодiї супротивника видiляють два випадки: супротивник вибирає послiдовнiсть винагород на початку горизонту, тобто вiн не займається вивченням стратегiї агента; супротивник обирає винагороди на кожному кроцi, що часто формулюють в термiнах теорiї iгор, а також використовують критерiй мiнiмаксу. В iншому варiантi моделi марковських бандитiв перехiд дiї у новий стан вiдбувається на кожному кроцi незалежно вiд того, вибрана ця дiя чи нi Метою моделi є пошук послiдовностi дiй, яка призводить до найбiльшої можливої сукупної винагороди

Модель стохастичного багаторукого бандита

Асимптотичний аналiз втрат

Стратегiї для стохастичного багаторукого бандита

Чисельнi експерименти

Зв’язок з марковським процесом прийняття рiшень

Висновки

Full Text

Published Version (Free)

View/Download pdf

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

Послiдовний розподiл ресурсiв устохастичному середовищi: загальнийопис, аналiз та чисельнi експерименти

Abstract

Highlights

Summary

Published Version (Free)

Talk to us

Similar Papers

More From: Bulletin of Taras Shevchenko National University of Kyiv. Series: Physics and Mathematics

Lead the way for us

Journal: Bulletin of Taras Shevchenko National University of Kyiv. Series: Physics and Mathematics	Publication Date: Jan 1, 2021
License type: cc-by

Similar Papers

Arm order recognition in multi-armed bandit problem with laser chaos time series
Naoki Narisawa ... Makoto Naruse
Scientific Reports | VOL. 11
Naoki Narisawa, et. al.Naoki Narisawa ... Makoto Naruse
24 Feb 2021
Scientific Reports | VOL. 11

Multi-armed bandit problem with online clustering as side information
Andrii Dzhoha ... Iryna Rozora
Journal of Computational and Applied Mathematics | VOL. 427
Andrii Dzhoha, et. al.Andrii Dzhoha ... Iryna Rozora
16 Feb 2023
Journal of Computational and Applied Mathematics | VOL. 427

Simultaneously Learning Stochastic and Adversarial Bandits under the Position-Based Model
Cheng Chen ... Canzhe Zhao
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 36
Cheng Chen, et. al.Cheng Chen ... Canzhe Zhao
28 Jun 2022
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 36

A Multi-armed Bandit Algorithm Available in Stationary or Non-stationary Environments Using Self-organizing Maps
Nobuhito Manome ... Kouta Suzuki
-
Nobuhito Manome, et. al.Nobuhito Manome ... Kouta Suzuki
01 Jan 2019
01 Jan 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

Послiдовний розподiл ресурсiв устохастичному середовищi: загальнийопис, аналiз та чисельнi експерименти

Abstract

Highlights

Summary

Published Version (Free)

Talk to us

Similar Papers

More From: Bulletin of Taras Shevchenko National University of Kyiv. Series: Physics and Mathematics