[an error occurred while processing this directive] [an error occurred while processing this directive]
[an error occurred while processing this directive]收稿日期: 2026-01-21
网络出版日期: 2026-05-09
基金资助
国家自然科学基金(62201587)
PPO-based Adaptive Frequency-hopping Anti-jamming Technique for Radar
Received date: 2026-01-21
Online published: 2026-05-09
针对现代复杂电磁环境中,处于稳定跟踪阶段的捷变频雷达在压制干扰下难以兼顾相干积累增益与抗干扰能力的问题,提出一种基于近端策略优化的频率与驻留时间联合控制方法。首先,分析了雷达相干积累增益与被截获风险之间的制约关系,模拟了干扰机的动态组合压制行为,构建了包含信号级特征反馈的雷达抗干扰马尔可夫决策过程模型,其次,引入近端策略优化算法,设计了兼顾探测效能与生存能力的复合奖励函数,使雷达智能体能够在未知的动态博弈环境中自主学习最优策略。该策略能够根据环境反馈实时调整工作频率与驻留时长,从而最大化信干噪比。仿真结果表明,与传统策略相比,所提方法能有效规避组合压制干扰,在显著降低被干扰概率的同时大幅提升了输出信干噪比,具有较强的环境适应性与鲁棒性。
杜钰渲 , 吴中杰 , 宋思盛 , 张军 . 基于PPO算法的雷达自适应跳频抗干扰技术研究[J]. 弹箭与制导学报, 2026 , 46(2) : 162 -171 . DOI: 10.15892/j.cnki.djzdxb.2026.02.005
To address the challenge that the frequency-agile radars in the stable tracking phase in modern complex electromagnetic environments struggle have difficulty in balancing both coherent integration gain and anti-jamming capability under blanket jamming,this paper proposes a joint control method for frequency and dwell time based on proximal policy optimization (PPO).Firstly,the trade-off relationship between radar coherent integration gain and intercept risk is analyzed.The dynamic combined jamming behavior of jammers is modeled,and a Markov decision process model for radar anti-jamming,incorporating signal-level feature feedback,is constructed.Secondly,the PPO algorithm is introduced.A compound reward function that takes into account both detection performance and survivability is designed,enabling the radar agent to autonomously learn the optimal strategy in unknown dynamical adversarial environments.This strategy can adjust the operating frequency and dwell time in real-time based on environmental feedback to maximize the signal-to-interference-plus-noise ratio.Simulated results demonstrate that,compared to traditional strategies,the proposed strategy can effectively evade the combined blanket jamming.It significantly reduces the probability of being jammed while substantially improving the output SINR,exhibiting strong environmental adaptability and robustness.
| [1] |
全英汇, 方文, 沙明辉, 等. 频率捷变雷达波形对抗技术现状与展望[J]. 系统工程与电子技术, 2021, 43(11):3126-3136.
|
| [2] |
全英汇, 方文, 高霞, 等. 捷变频雷达导引头技术现状与发展趋势[J]. 航空兵器, 2021, 28(3):1-9.
|
| [3] |
|
| [4] |
|
| [5] |
陈舒波. 基于深度强化学习的雷达抗有源干扰决策方法研究[D]. 成都: 电子科技大学, 2025.
|
| [6] |
刘全, 翟建伟, 章宗长, 等. 深度强化学习综述[J]. 计算机学报, 2018, 41(1):1-27.
|
| [7] |
李云杰, 朱云鹏, 高梅国. 基于Q-学习算法的认知雷达对抗过程设计[J]. 北京理工大学学报, 2015, 35(11):1194-1199.
|
| [8] |
邢强, 贾鑫, 朱卫纲. 基于Q-学习的智能雷达对抗[J]. 系统工程与电子技术, 2018, 40(5):1031-1035.
|
| [9] |
汪浩, 王峰. 强化学习算法在雷达智能抗干扰中的应用[J]. 现代雷达, 2020, 42(3):40-44.
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
解烽, 刘环宇, 胡锡坤, 等. 基于复数域深度强化学习的多干扰场景雷达抗干扰方法[J]. 雷达学报, 2023, 12(6):1290-1304.
|
| [22] |
李康, 纠博, 赵宇, 等. 雷达智能博弈抗干扰技术综述与展望[J]. 现代雷达, 2023, 45(5):15-26.
|
/
| 〈 |
|
〉 |