[an error occurred while processing this directive] [an error occurred while processing this directive]
[an error occurred while processing this directive]收稿日期: 2026-03-05
网络出版日期: 2026-06-29
VM-Net:A Distributed Multi-aircraft Mission Planning Method
Received date: 2026-03-05
Online published: 2026-06-29
针对复杂动态环境下多机任务规划中面临的长程信用分配(Long-term Credit Assignment Problem)和多机协同规划难题,提出一种基于价值混合网络(Value Mix Network,VM-Net)的分布式多机任务规划强化学习方法。首先,设计智能体级预测模块(VM-P),集成改进门控循环单元I-GRU与自注意力机制,实现在动作决策和状态评价中充分考虑历史轨迹;其次,设计值混合模块(VM-M),聚合个体价值函数为全局价值函数,在CTDE范式下实现智能体间显式协同;最后,设计几何感知奖励函数,有效缓解任务规划中的环境奖励稀疏问题。在对称(最高12v12)和非对称(6v24)场景下的大量实验表明,VM-Net优于传统的和基线强化学习方法。即使在其他方法性能下降的复杂环境中,VM-Net仍能保持实时推理(<20ms)和高胜率(>80%)。同时,训练过程中多智能体间合作行为的涌现进一步验证了VM-Net在动态任务规划方面的实际价值。
单永志 , 赵鹏博 , 李荣 , 郭建明 , 刘通 , 关秋雨 , 周增昊 . VM-Net:一种分布式多机任务规划算法[J]. 弹箭与制导学报, 2026 , 46(3) : 305 -317 . DOI: 10.15892/j.cnki.djzdxb.2026.03.008
To address the long-term credit assignment and multi-agent coordinative planning challenges in dynamic multi-aircraft mission planning,this paper proposes a novel reinforcement learning framework based on Value Mix Network (VM-Net).VM-Net consists of three key components,i.e.,agent-level prediction module (VM-P),value mixing module (VM-M) and geometry-aware dense reward function,working in concert under the centralized training and decentralized execution (CTDE) paradigm.The VM-P integrates an improved gated recurrent unit (I-GRU) with a self-attention mechanism,enabling the agents to leverage the historical trajectories for accurate state evaluation and policy learning.The VM-M aggregates individual value functions into a global Q-function to achieve the explicit cooperation among agents under the CTDE framework.The geometry-aware dense reward function based on relative distance and angle effectively mitigates the sparse reward problem in task planning and accelerates policy convergence.Extensive experiments in both symmetric (up to 12v12) and asymmetric (6v24) scenarios demonstrate that VM-Net outperforms SAC,MAPPO and QMIX.It is still able to maintain real-time inference (<20ms per step) and high win rates (>80%) even in complex environments where baselines degrade significantly.The emergence of sophisticated cooperative behaviors during training—such as coordinated flanking and dynamic target allocation—further validates VM-Net’s practical utility and its potential for deployment in real-world adversarial environments.
| [1] |
张文博, 刘君兰, 李瑜, 等. 多无人机任务与航迹规划方法综述[J]. 计算机与网络, 2022, 48(20):52-56.
|
| [2] |
唐瑞东, 游向荣. 基于改进人工势场法的多目标点路径规划[J]. 农业装备与车辆工程, 2023, 61(8):105-109,119.
|
| [3] |
丁祎男, 刘羽白, 王淑一, 等. 一种多目标变邻域模拟退火算法及成像星座任务规划方法[J]. 宇航学报, 2022, 43(12):1686-1695.
|
| [4] |
刘晓莹. 混沌蚁群算法在多机器人任务规划中的应用研究[D]. 长沙: 中南大学, 2010.
|
| [5] |
彭思文, 梁科, 潘雷, 等. 基于改进粒子群算法的机械臂时间最优轨迹规划[J]. 组合机床与自动化加工技术, 2023(6):31-34.
|
| [6] |
王超, 于德洋, 王子强, 等. 复杂任务环境下多无人机多任务规划技术研究[J]. 电子技术与软件工程, 2021(22):106-109.
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
王明君, 吴庆宪. 多无人机协同机场驱鸟任务分配[J]. 吉林大学学报(信息科学版), 2019, 37(1):47-57.
|
| [11] |
|
| [12] |
杜伟伟, 陈小伟. 陆军战术级作战任务分配及优化方法[J]. 兵工学报, 2023, 44(5):1431-1442.
|
| [13] |
封佳祥, 江坤颐, 周彬, 等. 多任务约束条件下基于强化学习的水面无人艇路径规划算法[J]. 舰船科学技术, 2019, 41(12):140-146.
|
| [14] |
施伟, 冯旸赫, 程光权, 等. 基于深度强化学习的多机协同空战方法研究[J]. 自动化学报, 2021, 47(7):1610-1623.
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
|
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
|
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
|
| [42] |
|
| [43] |
|
| [44] |
|
| [45] |
|
| [46] |
|
/
| 〈 |
|
〉 |