Multiagent Trust Region Policy Optimization.

Hepeng Li,Haibo He

doi:10.1109/tnnls.2023.3265358

Abstract

We extend trust region policy optimization (TRPO) to cooperative multiagent reinforcement learning (MARL) for partially observable Markov games (POMGs). We show that the policy update rule in TRPO can be equivalently transformed into a distributed consensus optimization for networked agents when the agents' observation is sufficient. By using a local convexification and trust-region method, we propose a fully decentralized MARL algorithm based on a distributed alternating direction method of multipliers (ADMM). During training, agents only share local policy ratios with neighbors via a peer-to-peer communication network. Compared with traditional centralized training methods in MARL, the proposed algorithm does not need a control center to collect global information, such as global state, collective reward, or shared policy and value network parameters. Experiments on two cooperative environments demonstrate the effectiveness of the proposed method.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Multiagent Trust Region Policy Optimization.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on neural networks and learning systems

Lead the way for us

Journal: IEEE transactions on neural networks and learning systems	Publication Date: Sep 1, 2024
Citations: 4

Similar Papers

Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs
Lior Shani ... Shie Mannor
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34
Lior Shani, et. al.Lior Shani ... Shie Mannor
03 Apr 2020
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34

Distributed policy evaluation via inexact ADMM in multi-agent reinforcement learning
Xiaoxiao Zhao ... Peng Yi
Control Theory and Technology | VOL. 18
Xiaoxiao Zhao, et. al.Xiaoxiao Zhao ... Peng Yi
25 Nov 2020
Control Theory and Technology | VOL. 18

Research on Supply Chain Optimization and Management Based on Deep Reinforcement Learning
Gao Yunxiang ... Wang Zhao
Scalable Computing: Practice and Experience | VOL. 25
Gao Yunxiang, et. al.Gao Yunxiang ... Wang Zhao
01 Oct 2024
Scalable Computing: Practice and Experience | VOL. 25

Authentic Boundary Proximal Policy Optimization.
Yuhu Cheng ... Longyang Huang
IEEE transactions on cybernetics | VOL. 52
Yuhu Cheng, et. al.Yuhu Cheng ... Longyang Huang
11 Mar 2021
IEEE transactions on cybernetics | VOL. 52

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Multiagent Trust Region Policy Optimization.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on neural networks and learning systems