[an error occurred while processing this directive] [an error occurred while processing this directive] [an error occurred while processing this directive]
[an error occurred while processing this directive]

基于强化学习的无人机智能任务分配方法

  • 费陈 ,
  • 郑晗 ,
  • 赵亮
展开
  • 武警士官学校基础部,杭州 311400

费陈(1994—),男,安徽芜湖人,硕士,研究方向:机器学习及强化学习、无人机应用。

收稿日期: 2022-09-24

  网络出版日期: 2025-02-25

Reinforcement Learning-based Intelligent Task Assignment Method for Unmanned Aerial Vehicles

  • FEI Chen ,
  • ZHENG Han ,
  • ZHAO Liang
Expand
  • Basic Department,Armed Police Officer School,Hangzhou 311400, China

Received date: 2022-09-24

  Online published: 2025-02-25

摘要

针对无人机群目标打击任务分配问题,提出一种基于强化学习的无人机智能任务分配方法。该方法提出一种任务分层框架,将多个无人机视为一个联盟并对目标进行分类,形成任务簇,并映射到无人机联盟中,通过多智能体强化学习算法(MADDPG)将任务簇内的目标与无人机联盟内的小无人机进行合理配对并对目标实施打击,得到MADDPG算法的回报值和飞行轨迹,并与DDPG算法、DQN算法的回报值和飞行轨迹进行对比。仿真结果表明,在小样本任务分配中,与不分层方法相比,该方法可以提高目标任务打击完成度,提升目标打击的效率;在分层框架下,相比于其他两种算法,收敛速度更快,收敛过程更加稳定。

本文引用格式

费陈 , 郑晗 , 赵亮 . 基于强化学习的无人机智能任务分配方法[J]. 弹箭与制导学报, 2022 , 42(6) : 61 -67 . DOI: 10.15892/j.cnki.djzdxb.2022.06.009

Abstract

Aiming at the task assignment problem of UAV swarm target strike, this paper proposes an intelligent UAV task assignment method based on reinforcement learning. This strategy proposes a task layering framework, which treats multiple UAVs as an alliance and classifies the targets to form task clusters, maps each task cluster to the UAV alliance. Through multi-agent reinforcement learning algorithm (MADDPG), the targets in the task cluster are reasonably paired with the small UAVs in the UAV alliance, then the targets are hit. The return value and flight path of MADDPG algorithm are obtained, and compared with the return value and flight path of DDPG algorithm and DQN algorithm. The experimental results show that in the task assignment of small samples, compared with the non-hierarchical method, this method can improve the completion degree of target task strike and improve the efficiency of target strike; under the hierarchical framework, compared with the other two algorithms, the convergence speed is faster, the convergence process is more stable.

[an error occurred while processing this directive]
[1]
武帅, 蔡春伟, 陈轶. 多旋翼无人机无线充电技术研究进展与发展趋势[J]. 电工技术学报, 2022, 37(3):555-565.

[2]
陈新颖, 盛敏. 面向6G的无人机通信综述[J]. 电子与信息学报, 2022, 44(3):781-789.

[3]
盛碧琦, 孙盛智, 侯妍. 美国无人机在局部作战行动中的运用及发展趋势[J]. 飞航导弹, 2020, 36(2):51-54.

[4]
刘晨. 一种侦察打击一体化无人机地面站软件的设计与实现[J]. 传感器世界, 2018, 24(7):31-37.

[5]
费思远, 鲜斌, 王岭. 基于群集行为的分布式多无人机编队动态避障控制[J]. 控制理论与应用, 2022, 39(1):1-11.

[6]
朱献文, 王端. 带有时延的多无人机编队同步控制的策略研究[J]. 组合机床与自动化加工技术, 2017(4):78-81.

DOI

[7]
ZHANG Q, WANG H, FENG Z, et al. Many-to-many matching-theory-based dynamic bandwidth allocation for UAVs[J]. IEEE Internet of Things Journal, 2021, 8(12): 9995-10009.

[8]
DING R, GAO F, SHEN X S. 3D UAV trajectory design and frequency band allocation for energy-efficient and fair communication: a deep reinforcement learning approach[J]. IEEE Transactions on Wireless Communications, 2020, 19(12): 7796-7809.

[9]
贾高伟, 王建峰. 无人机集群任务规划方法研究综述[J]. 系统工程与电子技术, 2021, 43(1):99-111.

DOI

[10]
HU J, ZHANG H, LIU L, et al. Convergent multiagent formation control with collision avoidance[J]. IEEE Transactions on Robotics, 2020, 36(6): 1805-1818.

[11]
刘畅, 谢文俊, 张鹏. 多目标群多基地多无人机协同任务规划[J]. 弹箭与制导学报, 2019, 39(1):119-124.

[12]
胡大伟, 陈海妹, 梁一为. 车辆与无人机混合编队的路径优化问题模型构建[J]. 长安大学学报(自然科学版), 2021, 41(1):78-89.

[13]
CHOI H L, BRUNET L, HOW J P. Consensus-based decentralized auctions for robust task allocation[J]. IEEE Transactions on Robotics, 2009, 25(4): 912-926.

[14]
RASMUSSEN S J, SHIMA T. Tree search algorithm for assigning cooperating UAVs to multiple tasks[J]. International Journal of Robust and Nonlinear Control, 2008, 18(2): 135-153.

[15]
蒋硕, 袁小平. 改进PSO算法在多无人机协同任务分配中的应用[J]. 计算机应用研究, 2019, 36(11):3344-3347.

[16]
常松, 贾子彦. 基于改进合同网算法的多无人机任务分配[J]. 物联网技术, 2020, 10(5):98-100.

[17]
KIM B H. Implementation of the centralized control system for drone training[J]. International Journal of Engineering & Technology, 2018, 7(23): 379-383.

[18]
韩本刚, 毛师彬. 多无人机协同编队飞行设计方法[J]. 兵工自动化, 2019, 38(12):1-13.

[19]
郑红星, 郭继峰. 一种未知动态环境下异构无人机集群分布式联盟形成方法[J]. 宇航学报, 2022, 43(2):189-197.

[20]
李波, 越凯强. 基于MADDPG的多无人机协同任务决策[J]. 宇航学报, 2021, 42(6):757-765.

[21]
符小卫, 潘静. 无人机集群规避动态障碍物的分布式队形控制[J]. 系统工程与电子技术, 2022, 44(2):529-537.

DOI

[22]
吴青坡, 周绍磊. 基于集散式模型预测控制的多无人机协同分区搜索[J]. 控制理论与应用, 2015, 32(10):1414-1421.

[23]
QIU H, DUAN H. Multiple UAV distributed close formation control based on in-flight leadership hierarchies of pigeon flocks[J]. Aerospace Science and Technology, 2017, 70(23): 471-486.

[24]
EPPE M, GUMBSCH C, KERZEL M, et al. Intelligent problem-solving as integrated hierarchical reinforcement learning[J]. Nature Machine Intelligence, 2022, 4(1): 11-20.

[25]
ZHANG Y, MOU Z, GAO F, et al. UAV-enabled secure communications by multi-agent deep reinforcement learning[J]. IEEE Transactions on Vehicular Technology, 2020, 69(10): 11599-11611.

文章导航

/

[an error occurred while processing this directive]