Official implementation of Heterogeneity in Multi-Agent Reinforcement Learning (AAMAS 2026).
Authors: Tianyi Hu, Zhiqiang Pu, Yuan Wang, Tenghai Qiu, Min Chen, Xin Yu
Affiliations: Institute of Automation, Chinese Academy of Sciences; National Key Laboratory of Cognition and Decision Intelligence for Complex Systems; University of Chinese Academy of Sciences
Heterogeneity is a fundamental property in multi-agent reinforcement learning (MARL), which is closely related not only to the functional differences of agents, but also to policy diversity and environmental interactions. However, the MARL field currently lacks a rigorous definition and deeper understanding of heterogeneity. This paper systematically discusses heterogeneity in MARL from the perspectives of definition, quantification, and utilization.
- Defining Heterogeneity: Based on an agent-level model of MARL, we categorize heterogeneity into five types: observation heterogeneity, response transition heterogeneity, effect transition heterogeneity, objective heterogeneity, and policy heterogeneity.
- Quantifying Heterogeneity: We define the heterogeneity distance and propose a quantification method based on representation learning (CVAE), applicable to both model-free and model-based settings. We also introduce Meta-Heterogeneity Distance to quantify agents' comprehensive heterogeneity.
- Utilizing Heterogeneity: We develop HetDPS (Heterogeneity-based Dynamic Parameter Sharing), a multi-agent algorithm that offers better interpretability and fewer task-specific hyperparameters compared to other parameter-sharing methods.
We recommend using a Conda virtual environment:
conda create --name hetdps python=3.8
conda activate hetdps
pip install -r requirements.txtNote: Some libraries may have dependency order issues. If automatic installation fails, try installing them manually.
- PettingZoo MPE (Particle-based Multi-agent Spreading): PettingZoo Version
- SMAC (StarCraft Multi-Agent Challenge): SMAC
Run HetDPS on the Particle-based Multi-agent Spreading environment:
python main_Het.py with algorithm_mode='hetdps' env_name='pettingzoo:pz-mpe-large-spread-v1' time_limit=50 parallel_envs=32 experiment_label='AAMAS'Run on different environment variants (v1–v4):
python main_Het.py with algorithm_mode='hetdps' env_name='pettingzoo:pz-mpe-large-spread-v2' time_limit=50 parallel_envs=32 experiment_label='AAMAS'
python main_Het.py with algorithm_mode='hetdps' env_name='pettingzoo:pz-mpe-large-spread-v3' time_limit=50 parallel_envs=32 experiment_label='AAMAS'
python main_Het.py with algorithm_mode='hetdps' env_name='pettingzoo:pz-mpe-large-spread-v4' time_limit=50 parallel_envs=32 experiment_label='AAMAS'| File | Description |
|---|---|
main_Het.py |
Main training loop: environment setup, sampling, PPO training, HetDPS scheduling, visualization |
model_Het.py |
Neural networks: MAPSNet, Policy, ConditionalLinearVAE, BiHierarchicalVAE |
utils_Het.py |
Heterogeneity quantification and HetDPS algorithm |
wrappers.py |
Environment wrappers for MARL benchmarks |
1. Heterogeneity Distance Computing (utils_Het.py)
compute_fusions— Policy heterogeneity distance (model-based): trains Conditional VAE, computes BD/Hellinger/WD between agentscompute_implicit_het— Meta-heterogeneity distance (model-free): trains BiHierarchical VAE, quantifies comprehensive agent heterogeneitycompute_dynamic_parameter_sharing— HetDPS: clustering via Affinity Propagation, network assignment with bipartite matchingcalculate_N_Gaussians_BD— Bhattacharyya distance (parallel, PyTorch)calculate_N_Gaussians_Hellinger_through_BD— Hellinger distancecalculate_N_Gaussians_WD— Wasserstein distance
2. Neural Network Models (model_Het.py)
MAPSNet— Multi-agent network with hierarchical parameter sharing (shallow + deep)Policy— Actor-Critic policy with dynamic parameter sharing supportConditionalLinearVAE— CVAE for policy distance (model-based)BiHierarchicalVAE— Hierarchical VAE for meta-heterogeneity (model-free)
3. Training Framework (main_Het.py)
- Multi-agent environment construction (PettingZoo, SMAC)
- Replay buffer and sample pool for heterogeneity measurement
- PPO + GAE training
- Periodic HetDPS execution and visualization
Our method achieves optimal or comparable performance across all tested tasks (PMS and SMAC) with identical hyperparameters, without task-specific tuning. Key findings:
- Interpretability: Meta-Het distance matrices align with agent type distributions and reveal emergent role division.
- Adaptability: HetDPS is robust to quantization intervals (20–2000 updates) and requires no cluster count or fusion threshold.
- Efficiency: Periodic heterogeneity quantification does not significantly reduce training speed.
Heterogeneity in Multi-Agent Reinforcement Learning (arXiv)
If you find this work useful, please cite:
@inproceedings{hu2026heterogeneity,
title={Heterogeneity in Multi-Agent Reinforcement Learning},
author={Hu, Tianyi and Pu, Zhiqiang and Wang, Yuan and Qiu, Tenghai and Chen, Min and Yu, Xin},
booktitle={Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)},
year={2026},
address={Paphos, Cyprus}
}This work was supported by the National Natural Science Foundation of China (Grants 62322316, 62503472), the Open Fund of National Key Laboratory of Information Systems Engineering (No. 6142101240203), and the Young Scientists Foundation of CSAA (GNC) under Grant CSAA-YSF2025-GNC-08.
- First Author: hutianyi2021@ia.ac.cn (Tianyi Hu)
- Corresponding Author: zhiqiang.pu@ia.ac.cn (Zhiqiang Pu)