Cooperative Multiagent Reinforcement Learning With Partial Observations

Yan Zhang,Michael M Zavlanos

doi:10.1109/tac.2023.3288025

Abstract

In this paper, we propose a distributed zeroth-order policy optimization method for Multi-Agent Reinforcement Learning (MARL). Existing MARL algorithms often assume that every agent can observe the states and actions of all the other agents in the network. This can be impractical in large-scale problems, where sharing the state and action information with multi-hop neighbors may incur significant communication overhead. The advantage of the proposed zeroth-order policy optimization method is that it allows the agents to compute the local policy gradients needed to update their local policy functions using local estimates of the global accumulated rewards that depend on partial state and action information only and can be obtained using consensus. Specifically, to calculate the local policy gradients, we develop a new distributed zeroth-order policy gradient estimator that relies on one-point residual-feedback which, compared to existing zeroth-order estimators that also rely on one-point feedback, significantly reduces the variance of the policy gradient estimates improving, in this way, the learning performance. We show that the proposed distributed zeroth-order policy optimization method with constant stepsize converges to the neighborhood of a policy that is a stationary point of the global objective function. The size of this neighborhood depends on the agents' learning rates, the exploration parameters, and the number of consensus steps used to calculate the local estimates of the global accumulated rewards. Moreover, we provide numerical experiments that demonstrate that our new zeroth-order policy gradient estimator is more sample-efficient compared to other existing one-point estimators.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Cooperative Multiagent Reinforcement Learning With Partial Observations

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Automatic Control

Lead the way for us

Similar Papers

Cooperative Multi-Agent Reinforcement Learning with Hierarchical Relation Graph under Partial Observability
Yang Li ... Xiangfeng Luo
-
Yang Li, et. al.Yang Li ... Xiangfeng Luo
01 Nov 2020
01 Nov 2020

Cooperative Multi-Agent Reinforcement Learning With Approximate Model Learning
Young Joon Park ... Seoung Bum Kim
IEEE Access | VOL. 8
Young Joon Park, et. al.Young Joon Park ... Seoung Bum Kim
01 Jan 2020
IEEE Access | VOL. 8

Communication-Free Two-Stage Multi-Agent DDPG under Partial States and Observations
Joohyun Cho ... Rong-Rong Chen
-
Joohyun Cho, et. al.Joohyun Cho ... Rong-Rong Chen
31 Oct 2021
31 Oct 2021

MDPGT: Momentum-Based Decentralized Policy Gradient Tracking
Zhanhong Jiang ... Sin Yong Tan
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 36
Zhanhong Jiang, et. al.Zhanhong Jiang ... Sin Yong Tan
28 Jun 2022
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 36

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Cooperative Multiagent Reinforcement Learning With Partial Observations

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Automatic Control