Unmanned Collective Intelligence Team Institute of Automation, Chinese Academy of Sciences (CASIA)
Curating high-quality open-source research code in Multi-Agent RL, LLM, and Robotics
We are the Unmanned Collective Intelligence Team (ๆ ไบบ้็พคๆบ่ฝๅข้), led by Prof. Zhiqiang Pu (่ฒๅฟๅผบ) at the Institute of Automation, Chinese Academy of Sciences (CASIA).
CASIA-Collect-AI is our open-source code collection platform, curating and maintaining high-quality research code from our team and affiliated researchers in MARL, LLM, and robotics.
Our group is affiliated with the National Key Laboratory of Cognition and Decision Intelligence for Complex Systems (่ฎค็ฅไธๅณ็ญๆบ่ฝๅ จๅฝ้็นๅฎ้ชๅฎค), where Prof. Pu serves as deputy director. We are also part of the School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS).
Our research operates at the intersection of:
- ๐ค Multi-Agent Reinforcement Learning (MARL) โ cooperative, heterogeneous, large-scale
- ๐ง LLM ร MARL โ RL-based LLM fine-tuning, alignment, and parameter-efficient methods
- โฝ Football / Sports AI โ data-driven tactic generation, VLM-guided policy alignment
- ๐ Autonomous Unmanned Systems โ UAV formation control, nonlinear flight control
| Direction | Description | Representative Work |
|---|---|---|
| MARL Foundations | Heterogeneity, parameter sharing, policy distance | MADPS (paper) (AAMAS'24 Oral), HetDPS (paper) (AAMAS'26 Oral) |
| Sparse Reward MARL | Lazy agents, responsibility diffusion, curriculum learning | LazyAgents (paper) (ICML'23) |
| LLM ร MARL | Fine-tuning LLMs with cooperative MARL, MoE PEFT | CORY (paper) (NeurIPS'24) |
| MARL Platforms | General simulation environments for MARL research | Unreal-MAP (paper) (AAAI'26 Oral) |
| Football AI | VLM-based reward shaping, generative tactic discovery | V-GEPF (paper) (AAAI'25), TacEleven (paper) |
| Hierarchical Agents | Nested LM agents for complex long-horizon tasks | agent-matrix |
| Paper | Venue | Repo |
|---|---|---|
| Unreal-MAP: Unreal-Engine-Based General Platform for Multi-Agent RL | AAAI 2026 (Oral) | โ |
| HetDPS: Heterogeneity in Multi-Agent Reinforcement Learning | AAMAS 2026 (Oral) | โ |
| Paper | Venue | Repo |
|---|---|---|
| TacEleven: Generative Tactic Discovery for Football Open Play | arXiv:2511.13326 | โ |
| V-GEPF: Vision-Based Generic Potential Function for Policy Alignment in MARL | AAAI 2025 | โ |
| CoMoE: Contrastive Representation for MoE in Parameter-Efficient Fine-tuning | EMNLP 2025 | โ |
| Cognition-Oriented Multiagent Reinforcement Learning | IEEE TNNLS | โ |
| A Policy Resonance Approach to Solve Responsibility Diffusion in MARL | IEEE TNNLS | โ |
| Self-Clustering Hierarchical MARL With Extensible Cooperation Graph | IEEE TETCI | โ |
| Hybrid Actor-Critic for Physically Heterogeneous MARL | IEEE TCDS | โ |
| Efficient Multitask RL via Task-Specific Action Correction | IEEE TCDS | โ |
| Paper | Venue | Repo |
|---|---|---|
| CORY: Coevolving with the Other You โ Fine-Tuning LLM with Sequential Cooperative MARL | NeurIPS 2024 | โ |
| MAPD: Measuring Policy Distance for Multi-Agent Reinforcement Learning | AAMAS 2024 (Oral) | โ |
| Orientation and Decision-Making for Soccer Based on Sports Analytics and AI | IEEE/CAA JAS | โ |
| Fuzzy Feedback MARL for Adversarial Dynamic Multiteam Competitions | IEEE TFS | โ |
| QFuture: Learning Future Expectation Cognition in MARL | IEEE TCDS | โ |
| Long-Term and Short-Term Opponent Intention Inference for Football MARL | IEEE TCDS | โ |
| Multiexperience-Assisted Efficient Multiagent Reinforcement Learning | IEEE TNNLS | โ |
| Paper | Venue | Repo |
|---|---|---|
| Lazy Agents: A New Perspective on Solving Sparse Reward in MARL | ICML 2023 | โ |
| Attention Enhanced Reinforcement Learning for Multi-Agent Cooperation | IEEE TNNLS | โ |
| Deep RL for Multiagent Formation Control With Collision Avoidance | IEEE TSMC:S | โ |
| Deep-RL-Based Multitarget Coverage With Connectivity Guaranteed | IEEE TII | โ |
| Automatic Curriculum Learning for Large-Scale Cooperative MARL | IEEE TETCI | โ |
| Cognition-Driven Multiagent Policy Learning for Promoting Cooperation | IEEE TG | โ |
| Learning to Play Football From Sports Domain Perspective | IEEE TG | โ |
| Paper | Venue |
|---|---|
| ConcNet: Concentration Network for RL of Large-Scale Multi-Agent Systems | AAAI 2022 |
| Multi-Target Encirclement with Collision Avoidance via Deep RL | ICRA 2022 |
| Relative Distributed Formation and Obstacle Avoidance with Multi-Agent RL | ICRA 2022 |
| Fixed-Time Adaptive Fuzzy Control for Uncertain Nonstrict-Feedback Systems (Highly Cited, WoS) | IEEE TFS |
| Repository | Topic | Paper |
|---|---|---|
| LLM-MARL-CORY | LLM fine-tuning via cooperative MARL | NeurIPS 2024 |
| LLM-Football-TacEleven | LLM-based generative football tactic discovery | arXiv 2025 |
| MARL-Diversity-HetDPS | Heterogeneity & dynamic parameter sharing | AAMAS 2026 (Oral) |
| MARL-Diversity-MADPS | Policy distance metric for MARL | AAMAS 2024 (Oral) |
| MARL-Environment-UnrealMAP | Unreal Engine MARL simulation platform | AAAI 2026 (Oral) |
| MARL-Football-VGEPF | VLM-based reward shaping for football MARL | AAAI 2025 |
| MARL-Reward-LazyAgents | Sparse reward & lazy agent problem in MARL | ICML 2023 |
| agent-matrix | Hierarchical nested LM agents (new project) | โ |
๐ฅ ZKDX Intelligence (ไธญ็งๅบๆบ่ฝ) is hiring!
Embodied AI ยท Multimodal LLMs ยท Robotics
๐ View details: recruitment.md
- Prof. Zhiqiang Pu: zhiqiang.pu@ia.ac.cn
- Lab: National Key Laboratory of Cognition and Decision Intelligence for Complex Systems, CASIA, Beijing
- UCAS Profile: people.ucas.edu.cn/~pzq
