[an error occurred while processing this directive] [an error occurred while processing this directive] [an error occurred while processing this directive]
[an error occurred while processing this directive]

VM-Net:一种分布式多机任务规划算法

  • 单永志 1 ,
  • 赵鹏博 , 1, * ,
  • 李荣 2 ,
  • 郭建明 1 ,
  • 刘通 1 ,
  • 关秋雨 1 ,
  • 周增昊 1
展开
  • 1 中国兵器工业集团航空弹药研究院有限公司, 黑龙江 哈尔滨 150036
  • 2 中国人民解放军 93128 部队, 北京 100038

收稿日期: 2026-03-05

  网络出版日期: 2026-06-29

VM-Net:A Distributed Multi-aircraft Mission Planning Method

  • SHAN Yongzhi 1 ,
  • ZHAO Pengbo , 1, * ,
  • LI Rong 2 ,
  • GUO Jianming 1 ,
  • LIU Tong 1 ,
  • GUAN Qiuyu 1 ,
  • ZHOU Zenghao 1
Expand
  • 1 Air Ammunition Research Institute CO., LTD, Harbin 150036,Heilongjiang, China
  • 2 93128 Unit of the PLA No., Beijing 100038, China

Received date: 2026-03-05

  Online published: 2026-06-29

摘要

针对复杂动态环境下多机任务规划中面临的长程信用分配(Long-term Credit Assignment Problem)和多机协同规划难题,提出一种基于价值混合网络(Value Mix Network,VM-Net)的分布式多机任务规划强化学习方法。首先,设计智能体级预测模块(VM-P),集成改进门控循环单元I-GRU与自注意力机制,实现在动作决策和状态评价中充分考虑历史轨迹;其次,设计值混合模块(VM-M),聚合个体价值函数为全局价值函数,在CTDE范式下实现智能体间显式协同;最后,设计几何感知奖励函数,有效缓解任务规划中的环境奖励稀疏问题。在对称(最高12v12)和非对称(6v24)场景下的大量实验表明,VM-Net优于传统的和基线强化学习方法。即使在其他方法性能下降的复杂环境中,VM-Net仍能保持实时推理(<20ms)和高胜率(>80%)。同时,训练过程中多智能体间合作行为的涌现进一步验证了VM-Net在动态任务规划方面的实际价值。

本文引用格式

单永志 , 赵鹏博 , 李荣 , 郭建明 , 刘通 , 关秋雨 , 周增昊 . VM-Net:一种分布式多机任务规划算法[J]. 弹箭与制导学报, 2026 , 46(3) : 305 -317 . DOI: 10.15892/j.cnki.djzdxb.2026.03.008

Abstract

To address the long-term credit assignment and multi-agent coordinative planning challenges in dynamic multi-aircraft mission planning,this paper proposes a novel reinforcement learning framework based on Value Mix Network (VM-Net).VM-Net consists of three key components,i.e.,agent-level prediction module (VM-P),value mixing module (VM-M) and geometry-aware dense reward function,working in concert under the centralized training and decentralized execution (CTDE) paradigm.The VM-P integrates an improved gated recurrent unit (I-GRU) with a self-attention mechanism,enabling the agents to leverage the historical trajectories for accurate state evaluation and policy learning.The VM-M aggregates individual value functions into a global Q-function to achieve the explicit cooperation among agents under the CTDE framework.The geometry-aware dense reward function based on relative distance and angle effectively mitigates the sparse reward problem in task planning and accelerates policy convergence.Extensive experiments in both symmetric (up to 12v12) and asymmetric (6v24) scenarios demonstrate that VM-Net outperforms SAC,MAPPO and QMIX.It is still able to maintain real-time inference (<20ms per step) and high win rates (>80%) even in complex environments where baselines degrade significantly.The emergence of sophisticated cooperative behaviors during training—such as coordinated flanking and dynamic target allocation—further validates VM-Net’s practical utility and its potential for deployment in real-world adversarial environments.

[an error occurred while processing this directive]
[1]
张文博, 刘君兰, 李瑜, 等. 多无人机任务与航迹规划方法综述[J]. 计算机与网络, 2022, 48(20):52-56.

ZHANG W B, LIU J L, LI Y, et al. Review of multi-UAV mission and track planning methods[J]. Computer & Network, 2022, 48(20):52-56.

[2]
唐瑞东, 游向荣. 基于改进人工势场法的多目标点路径规划[J]. 农业装备与车辆工程, 2023, 61(8):105-109,119.

TANG R D, YOU X R. Multi-objective point path planning based on improved artificial potential field method[J]. Agricultural Equipment & Vehicle Engineering, 2023, 61(8):105-109,119.

[3]
丁祎男, 刘羽白, 王淑一, 等. 一种多目标变邻域模拟退火算法及成像星座任务规划方法[J]. 宇航学报, 2022, 43(12):1686-1695.

DING Y N, LIU Y B, WANG S Y, et al. A multi-objective variable-neighborhood simulated annealing algorithm and imaging constellation task planning method[J]. Journal of Astronautics, 2022, 43(12):1686-1695.

[4]
刘晓莹. 混沌蚁群算法在多机器人任务规划中的应用研究[D]. 长沙: 中南大学, 2010.

LIU X Y. Application of chaos ant colony algorithm in multi-robots system mission planning[D]. Changsha: Central South University, 2010.

[5]
彭思文, 梁科, 潘雷, 等. 基于改进粒子群算法的机械臂时间最优轨迹规划[J]. 组合机床与自动化加工技术, 2023(6):31-34.

PENG S W, LIANG K, PAN L, et al. Time optimal trajectory planning of manipulator based on improved particle swarm optimization algorithm[J]. Modular Machine Tool & Automatic Manufacturing Technique, 2023(6):31-34.

[6]
王超, 于德洋, 王子强, 等. 复杂任务环境下多无人机多任务规划技术研究[J]. 电子技术与软件工程, 2021(22):106-109.

WANG C, YU D Y, WANG Z Q, et al. Research on multi-mission planning technology of multi-UAV in complex mission environment[J]. Electronic Technology & Software Engineering, 2021(22):106-109.

[7]
AGIS R A, GOTTIFREDI S, GARCÍA A J. An event-driven behavior trees extension to facilitate non-player multi-agent coordination in video games[J]. Expert Systems with Applications, 2020,155:W0178A.

[8]
DAHLQUIST N, LINDQVIST B, SARADAGI A, et al. Reactive multi-agent coordination using auction-based task allocation and behavior trees[EB/OL].(2023-04-04)[2026-02-28]. https://arxiv.org/abs/2304.01976.

[9]
NEUFELD X, MOSTAGHIM S, BRAND S. A hybrid approach to planning and execution in dynamic environments through hierarchical task networks and behavior trees[C]//Proceedings of the Fourteenth AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment.Edmonton,NA,CAN:AAAI Press,2018:201-207.

[10]
王明君, 吴庆宪. 多无人机协同机场驱鸟任务分配[J]. 吉林大学学报(信息科学版), 2019, 37(1):47-57.

WANG M J, WU Q X. Cooperative task assignment of multi-UAVs bird-driving at airport[J]. Journal of Jilin University(Information Science Edition), 2019, 37(1):47-57.

[11]
TAN W, HU Y J, ZHAO Y F, et al. Multi-UAV task allocation based on type mamdani fuzzy logic[C]//Proceedings of the 2021 7th International Symposium on Mechatronics and Industrial Informatics (ISMII).Zhuhai,CN:IEEE,2021:184-187.

[12]
杜伟伟, 陈小伟. 陆军战术级作战任务分配及优化方法[J]. 兵工学报, 2023, 44(5):1431-1442.

DU W W, CHEN X W. Task assignment and optimization method of tactical-level army operations[J]. Acta Armamentarii, 2023, 44(5):1431-1442.

[13]
封佳祥, 江坤颐, 周彬, 等. 多任务约束条件下基于强化学习的水面无人艇路径规划算法[J]. 舰船科学技术, 2019, 41(12):140-146.

FENG J X, JIANG K Y, ZHOU B, et al. Path planning for USV based on reinforcement learning with multi-task constraints[J]. Ship Science and Technology, 2019, 41(12):140-146.

[14]
施伟, 冯旸赫, 程光权, 等. 基于深度强化学习的多机协同空战方法研究[J]. 自动化学报, 2021, 47(7):1610-1623.

SHI W, FENG Y H, CHENG G Q, et al. Research on multi-aircraft cooperative air combat method based on deep reinforcement learning[J]. Acta Automatica Sinica, 2021, 47(7):1610-1623.

[15]
BAYERLEIN H, THEILE M, CACCAMO M, et al. Multi-UAV path planning for wireless data harvesting with deep reinforcement learning[J]. IEEE Open Journal of the Communications Society, 2021, 2:1171-1187.

[16]
CHEN Y, DONG Q, SHANG X Z, et al. Multi-UAV autonomous path planning in reconnaissance missions considering incomplete information:a reinforcement learning method[J]. Drones, 2022, 7(1):10.

[17]
BEARD R W, MCLAIN T W, GOODRICH M A, et al. Coordinated target assignment and intercept for unmanned air vehicles[J]. IEEE Transactions on Robotics and Automation, 2022, 18(6):911-922.

[18]
BELLINGHAM J, RICHARDS A, How J P. Receding horizon control of autonomous aerial vehicles[C]//Proceedings of the 2002 American Control Conference (IEEE Cat.No.CH37301).Anchorage,AK,USA:IEEE,2002:3741-3746.

[19]
CASTANON D A, WOHLETZ J M. Model predictive control for stochastic resource allocation[J]. IEEE Transactions on Automatic Control, 2009, 54(8):1739-1750.

[20]
KIM Y, GU D W, POSTLETHWAITE I. Real-time optimal mission scheduling and flight path selection[J]. IEEE Transactions on Automatic Control, 2007, 52(6):1119-1123.

[21]
RASMUSSEN S J, SHIMA T. Tree search algorithm for assigning cooperating UAVs to multiple tasks[J]. International Journal of Robust and Nonlinear Control, 2008, 18(2):135-153.

[22]
LEARY S, DEITTERT M, BOOKLESS J. Constrained UAV mission planning:a comparison of approaches[C]//Proceedings of the 2011 IEEE International Conference on Computer Vision Workshops (ICCV Workshops).Barcelona,ESP:IEEE,2011:2002-2009.

[23]
SHIMA T, RASMUSSEN S, RASMUSSEN S J. UAV cooperative decision and control:challenges and practical approaches[M]. Philadelphia,PA, USA:SIAM, 2008.

[24]
XUE K, HUANG Z Q, WANG P, et al. An exact algorithm for task allocation of multiple unmanned surface vehicles with minimum task time[J]. Journal of Marine Science and Engineering, 2021, 9(8):907.

[25]
YE X N, LEI Z K, LIU K. Multi-UAV task assignment based on satisficing decision algorithm[C]//Proceedings of the 2021 IEEE International Conference on Consumer Electronics and Computer Engineering (ICCECE). Guangzhou,CN:IEEE,2021:439-442.

[26]
EDISON E, SHIMA T. Integrated task assignment and path optimization for cooperating uninhabited aerial vehicles using genetic algorithms[J]. Computers & Operations Research, 2011, 38(1):340-356.

[27]
SCHWARZROCK J, ZACARIAS I, BAZZAN A L C. Solving task allocation problem in multi unmanned aerial vehicles systems using swarm intelligence[J]. Engineering Applications of Artificial Intelligence, 2018, 72:10-20.

[28]
WU Y, LIANG T J, GOU J Z, et al. Heterogeneous mission planning for multiple UAV formations via metaheuristic algorithms[J]. IEEE Transactions on Aerospace and Electronic Systems, 2023, 59(4):3924-3940.

[29]
ZHEN Z Y, CHEN Y, WEN L D, et al. An intelligent cooperative mission planning scheme of UAV swarm in uncertain dynamic environment[J]. Aerospace Science and Technology, 2020, 100:105826.

[30]
LIU Y B, QI N M, YAO W R. Optimal scheduling for aerial recovery of multiple unmanned aerial vehicles using genetic algorithm[J]. Proceedings of the Institution of Mechanical Engineers, 2019, 233(14):5347-5359.

[31]
KIM J, OH H, YU B, et al. Optimal task assignment for UAV swarm operations in hostile environments[J]. International Journal of Aeronautical and Space Sciences, 2021, 22(2):456-467.

[32]
MNIH V, KAVUKCUOGLU K, SILVER D, et al. Human-level control through deep reinforcement learning[J]. Nature, 2015, 518(7540):529-533.

[33]
VAN HASSELT H, GUEZ A, SILVER D. Deep reinforcement learning with double Q-Learning[C]//Proceedings of the Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence. Phoenix,AZ: AAAI Press,2016:2094-2100.

[34]
SILVER D, LEVER G, HEESS N, et al. Deterministic policy gradient algorithms[C]//Proceedings of the 31st International Conference on Machine Learning. Beijing:JMLR,2014:I-387-I-395.

[35]
CHRISTODOULOU P. Soft actor-critic for discrete action settings[EB/OL].(2019-10-18)[2026-02-28].https://doi.org/10.48550/arXiv.1910.07207.

[36]
XIONG J C, WANG Q, YANG Z R, et al. Parametrized deep q-networks learning:Reinforcement learning with discrete-continuous hybrid action space[EB/OL].(2018-10-10)[2026-03-01].https://doi.org/10.48550/arXiv.1810.06394.

[37]
YU C, VELU A, VINITSKY E, et al. The surprising effectiveness of PPO in cooperative,multi-agent games[EB/OL].(2022-11-04)[2026-03-01].https://doi.org/10.48550/arXiv.2103.01955.

[38]
BELLO I, PHAM H, LE Q V, et al. Neural combinatorial optimization with reinforcement learning[EB/OL].(2017-07-12)[2026-03-01].https://doi.org/10.48550/arXiv.1611.09940.

[39]
KHALIL E, DAI H J, ZHANG Y Y, et al. Learning combinatorial optimization algorithms over graphs[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. Red Hook,NY,USA: Curran Associates,2017:6351-6361.

[40]
MOHAMMADREZA N, AFSHIN O, TAKÁČM, et al. Reinforcement learning for solving the vehicle routing problem[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. Red Hook,NY,USA: Curran Associates,2018:6351-6361.

[41]
DEUDON M, COURNUT P, LACOSTE A, et al. Learning heuristics for the TSP by policy gradient[C]//VAN HOEVE W J.Integration of Constraint Programming,Artificial Intelligence,and Operations Research. Cham,ZUG,CHE:Springer,2018:170-181.

[42]
DRORI I, KHARKAR A, SICKINGER W R, et al. Learning to solve combinatorial optimization problems on real-world graphs in linear time[C]//Proceedings of the 2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA).Miami,FL,USA:IEEE,2020:19-24.

[43]
KOOL W, VAN HOOF H, MAX W. Attention,learn to solve routing problems![EB/OL].(2019-02-07)[2026-03-04].https://doi.org/10.48550/arXiv.1803.08475.

[44]
XIN L, SONG W, CAO Z G, et al. Step-wise deep learning models for solving routing problems[J]. IEEE Transactions on Industrial Informatics, 2021, 17(7):4861-4871.

[45]
RASHID T, SAMVELYAN M, DE WITT C S, et al. Monotonic value function factorisation for deep multi-agent reinforcement learning[J]. The Journal of Machine Learning Research, 2020, 21(1):7234-7284.

[46]
KINGMA D P, BA J. Adam:a method for stochastic optimization[EB/OL].(2017-01-30)[2026-03-05].https://doi.org/10.48550/arXiv.1412.6980.

文章导航

/

[an error occurred while processing this directive]