Partner-Aware Algorithms in Decentralized Cooperative Bandit Teams

Erdem Biyik,Anusha Lalitha,Andrea Goldsmith,Dorsa Sadigh,Rajarshi Saha

doi:10.1609/aaai.v36i9.21158

Abstract

When humans collaborate with each other, they often make decisions by observing others and considering the consequences that their actions may have on the entire team, instead of greedily doing what is best for just themselves. We would like our AI agents to effectively collaborate in a similar way by capturing a model of their partners. In this work, we propose and analyze a decentralized Multi-Armed Bandit (MAB) problem with coupled rewards as an abstraction of more general multi-agent collaboration. We demonstrate that naive extensions of single-agent optimal MAB algorithms fail when applied for decentralized bandit teams. Instead, we propose a Partner-Aware strategy for joint sequential decision-making that extends the well-known single-agent Upper Confidence Bound algorithm. We analytically show that our proposed strategy achieves logarithmic regret, and provide extensive experiments involving human-AI and human-robot collaboration to validate our theoretical findings. Our results show that the proposed partner-aware strategy outperforms other known methods, and our human subject studies suggest humans prefer to collaborate with AI agents implementing our partner-aware strategy.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Partner-Aware Algorithms in Decentralized Cooperative Bandit Teams

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: Jun 28, 2022
Citations: 1

Similar Papers

Investigation of selection and application of Multi-Armed Bandit algorithms in recommendation system
Panyangjie Chen
Applied and Computational Engineering | VOL. 34
Panyangjie ChenPanyangjie Chen
04 Feb 2024
Applied and Computational Engineering | VOL. 34

A Multi-armed Bandit Algorithm Available in Stationary or Non-stationary Environments Using Self-organizing Maps
Nobuhito Manome ... Shuji Shinohara
-
Nobuhito Manome, et. al.Nobuhito Manome ... Shuji Shinohara
01 Jan 2019
01 Jan 2019

Statistical Consequences of using Multi-armed Bandits to Conduct Adaptive Educational Experiments
...
-
, et. al. ...
18 Jun 2019
18 Jun 2019

Adaptive operator selection with dynamic multi-armed bandits
Luis Dacosta ... Alvaro Fialho
-
Luis Dacosta, et. al.Luis Dacosta ... Alvaro Fialho
12 Jul 2008
12 Jul 2008

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Partner-Aware Algorithms in Decentralized Cooperative Bandit Teams

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence