Thompson Sampling for Stochastic Bandits with Noisy Contexts: An Information-Theoretic Regret Analysis.

Sharu Theresa Jose,Shana Moothedath

doi:10.3390/e26070606

Abstract

We study stochastic linear contextual bandits (CB) where the agent observes a noisy version of the true context through a noise channel with unknown channel parameters. Our objective is to design an action policy that can "approximate" that of a Bayesian oracle that has access to the reward model and the noise channel parameter. We introduce a modified Thompson sampling algorithm and analyze its Bayesian cumulative regret with respect to the oracle action policy via information-theoretic tools. For Gaussian bandits with Gaussian context noise, our information-theoretic analysis shows that under certain conditions on the prior variance, the Bayesian cumulative regret scales as O˜(mT), where m is the dimension of the feature vector and T is the time horizon. We also consider the problem setting where the agent observes the true context with some delay after receiving the reward, and show that delayed true contexts lead to lower regret. Finally, we empirically demonstrate the performance of the proposed algorithms against baselines.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Thompson Sampling for Stochastic Bandits with Noisy Contexts: An Information-Theoretic Regret Analysis.

Abstract

Talk to us

Similar Papers

More From: Entropy (Basel, Switzerland)

Lead the way for us

Journal: Entropy (Basel, Switzerland)	Publication Date: Jul 17, 2024
License type: CC BY 4.0

Similar Papers

Data-Aided SNR Estimation in Time-Variant Rayleigh Fading Channels
Habti Abeida
IEEE Transactions on Signal Processing | VOL. 58
Habti AbeidaHabti Abeida
01 Nov 2010
IEEE Transactions on Signal Processing | VOL. 58

Learning the unknown: Improving modulation classification performance in unseen scenarios
Erma Perenda ... Mariya Zheleva
-
Erma Perenda, et. al.Erma Perenda ... Mariya Zheleva
10 May 2021
10 May 2021

A Deep Autoencoder Approach to Received Signal Strength-Based Localization with Unknown Channel Parameters
Chaehun Im ... Chungyoung Lee
-
Chaehun Im, et. al.Chaehun Im ... Chungyoung Lee
01 Feb 2020
01 Feb 2020

Channel Estimation for Macrocellular OFDM Uplinks in Time-Varying Channels
Yinsheng Liu ... Kyung Sup Kwak
IEEE Transactions on Vehicular Technology | VOL. 61
Yinsheng Liu, et. al.Yinsheng Liu ... Kyung Sup Kwak
01 May 2012
IEEE Transactions on Vehicular Technology | VOL. 61

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Thompson Sampling for Stochastic Bandits with Noisy Contexts: An Information-Theoretic Regret Analysis.

Abstract

Talk to us

Similar Papers

More From: Entropy (Basel, Switzerland)