[an error occurred while processing this directive] [an error occurred while processing this directive] [an error occurred while processing this directive]
[an error occurred while processing this directive]

基于弹载图像的Transformer目标跟踪算法

  • 孙子文 1 ,
  • 钱立志 1 ,
  • 袁广林 2 ,
  • 凌冲 1
展开
  • 1 陆军炮兵防空兵学院高过载弹药制导控制与信息感知实验室,安徽 合肥 230031
  • 2 陆军炮兵防空兵学院信息工程系,安徽 合肥 230031
凌冲(1983—),男,讲师,硕士,研究方向:深度学习、图像处理系统开发。

孙子文(1996—),男,博士研究生,主要研究方向:计算机视觉跟踪。

收稿日期: 2023-09-19

  网络出版日期: 2024-12-28

Transformer Object Tracking Algorithm Based on Missile-borne Image

  • SUN Ziwen 1 ,
  • QIAN Lizhi 1 ,
  • YUAN Guanglin 2 ,
  • LING Chong 1
Expand
  • 1 High Overload Ammunition Guidance Control and Information Perception Laboratory of the Army Artillery Air Defense Academy, Hefei 230031, Anhui, China
  • 2 Department of Information Engineering, Army Academy of Artillery and Air Defense, Hefei 230031, Anhui, China

Received date: 2023-09-19

  Online published: 2024-12-28

摘要

基于弹载图像自寻的经典目标跟踪方法依靠传统特征提取方法通常只能提取到有限的特征,表征能力不足,算法易受弹载图像中目标尺度变化、相似目标以及复杂背景等因素干扰而导致跟踪失效。由于Transformer凭借强大的全局建模能力被广泛应用于目标跟踪领域。结合弹载图像仿真实验平台提出基于弹载图像的Transformer目标跟踪算法,算法由特征提取、特征融合以及预测头三部分组成。首先,在特征提取部分分别使用Swin-Transformer网络的前三层对输入的初始模板和搜索区域提取深度特征。其次,为了充分利用初始模板信息,借助交叉注意力模块对提取的特征进行特征增强处理。然后,将提取后的特征进行拼接并送入编码器、解码器模块进行特征的融合。最后,输出的特征经过回归和分类头进行目标定位。算法在弹载图像数据集上进行实验,跟踪成功率达到73.87%,跟踪速度达到56.79 帧/s。相较于经典的KCF算法,文中算法充分利用Transformer注意力机制的特性在跟踪成功率以及精度上提高了18.01%和23.14%,大幅提升算法的鲁棒性。

本文引用格式

孙子文 , 钱立志 , 袁广林 , 凌冲 . 基于弹载图像的Transformer目标跟踪算法[J]. 弹箭与制导学报, 2024 , 44(1) : 49 -56 . DOI: 10.15892/j.cnki.djzdxb.2024.01.008

Abstract

The classic object tracking method based on missile borne image self seeking relies on traditional feature extraction methods, which usually can only extract limited features and have insufficient representation ability, and the algorithms are susceptible to tracking failures due to the interference of factors such as changes in target scales, similar targets, and complex backgrounds in missile-borne images. Since Transformer is widely used in the field of object tracking by virtue of its powerful global modeling capability. In this paper, the Transformer object tracking algorithm based on missile-borne image is proposed by combining the experimental platform of missile-borne image simulation, which consists of three parts: feature extraction, feature fusion and prediction head. First, deep features are extracted from the input initial template and the search region using the first three layers of the Swin-Transformer network, respectively, in the feature extraction part. Second, in order to fully utilize the initial template information, the extracted features are feature enhanced with the help of the cross-attention module. Then, the extracted features are spliced and fed to the encoder and decoder modules for fusion of features. Finally, the output features are regressed and classified header for target localization. The algorithm is experimented on the missile-borne image dataset, and the tracking success rate reaches 73.87%, and the tracking speed reaches 56.79 frames/s. Compared to the classic KCF algorithm, the algorithm in this article fully utilizes the characteristics of the Transformer attention mechanism to improve the tracking success rate and accuracy by 18.01% and 23.14%, significantly enhancing the robustness of the algorithm.

[an error occurred while processing this directive]
[1]
VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[C]// NIIRS. Proceeding of the 31st International Conference on Neural Information Processing Systems. New York: ACM, 2017: 5998-6008.

[2]
YU B, TANG M, ZHENG L Y, et al. High-performance discriminative tracking with transformers[C]// IEEE. Proceedings of the 20 IEEE International Conference on Computer Vision. Montreal. New York: IEEE, 2021: 9836-9845.

[3]
CHEN X, YAN B, ZHU J W, et al. Transformer tracking[C]// CVPR. Proceedings of the 2021 Conference on Computer Vision and Pattern Recognition. New York: CVPR, 2021: 8126-8135.

[4]
WANG N, ZHOU W G, WANG J, et al. Transformer meets tracker: exploiting temporal context for robust visual tracking[C]// CVPR. Proceedings of the 2021 Conference on Computer Vision and Pattern Recognition. New York: CVPR, 2021: 1571-1580.

[5]
BLATTER P, KANAKIS M, DANELLJAN M, et al. Efficient visual tracking with exemplar transformers[C]// CVPR. Proceedings of the 2022 Winter Conference on Applications of Computer Vision. New York, 2022: 1571-1581.

[6]
SONG Z K, YU J Q, CHEN Y P, et al. Transformer tracking with cyclic shifting window attention[C]// CVPR. Proceedings of the 2022 Conference on Computer Vision and Pattern Recognition. New York: CVPR, 2022: 8781-8790.

[7]
MAYER C, DANELLJIA M, BHAT G, et al. Transforming model prediction for tracking[C]// CVPR. Proceedings of the 2022 Conference on Computer Vision and Pattern Recognition. New York: CVPR, 2022: 8721-8730.

[8]
GAO S Y, ZHOU C L, MA C. AiATrack: attention in attention for transformer visual tracking[C]// ECCV. Proceedings of the 2022 European Conference on Computer Vision. New York: IEEE, 2022: 146-164.

[9]
LIU Z, LIN Y, CAO Y, et al. Swin transformer: hierarchical vision transformer using shifted windows[C]// ICCV. Proceedings of the 2021 International Conference on Computer Vision. New York: IEEE, 2021: 9992-10002.

[10]
LIN L T, FAN H, XU Y, et al. SwinTrack: A simple and strong baseline for Transformer tracking[EB/OL]. https://doi.org/10.48550/arXiv.20112.00995.

[11]
TANG C M, WANG X, BAI Y C, et al. Learning spatial-frequency Transformer for visual object tracking[EB/OL]. https://doi.org/CoRR 2022.abs/2208.08829.

[12]
QIAO J F, LI F J, HAN H G, et al. Constructive algorithm for fully connected cascade feedforward neural networks[J]. Neurocomputing, 2016(182): 154-164.

[13]
HE K M, ZHANG X Y, REN S Q, et al. Deep residual learning for image recognition[C]// CVPR. Proceedings of the 2016 Conference on Computer Vision and Pattern Recognition. New York: CVPR, 2016: 770-778.

[14]
KRIZHEVSKY A, SUTSKEVER I, E. H GEOFFREY. ImageNet classification with deep convolutional neural networks[J]. Communication. ACM, 2017, 60(6): 84-90.

[15]
PETER S, JAKOB U, ASHISH V. Self-attention with relative position representations[C]//IEEE. Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies, NAACL-HLT. New York: IEEE, 2018: 464-468.

[16]
EDOUARD G, ARMAND J, MOUSTAPHA C, et al. Efficient softmax approximation for GPUs[C]// ICML. Proceedings of the 2017 International Conference on Machine Learning. New York: ACM, 2017: 1302-1310.

[17]
VINOD N, GEOFFREY E H. Rectified linear units improve restricted boltzmann machines[C]// ICML. Proceedings of the 2010 International Conference on Machine Learning. New York: ACM, 2010: 807-814.

[18]
FAN H, LING H, LIN L, et al. LaSOT: a high quality benchmark for large scale single object tracking[C]// CVPR. Proceedings of the 2019 Conference on Computer Vision and Pattern Recognition. New York: CVPR, 2019: 5374-5383.

[19]
MULLER M, BIBI A, GIANCOLA S, et al. TrackingNet: a large scale dataset and benchmark for object tracking in the wild[C]// ECCV. Proceedings of the 2018 European Conference on Computer Vision. New York: IEEE, 2018: 300-317.

[20]
HUANG L, ZHAO X, HUANG K. GOT-10k: a large high diversity benchmark for generic object tracking in the wild[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43(5): 1562-1577.

[21]
LOSHCHILOV I, HUTTER F. Decoupled weight decay regularization[C]// ICML. Proceedings of the 2019 International Conference on Learning Representations. New York: ACM, 2019: 1-18.

[22]
HENRIQUES J F, CASEIRO R, MARTINS P, et al. High-speed tracking with kernelized correlation filters[J]. IEEE Trans Pattern Anal Mach Intell, 2015, 37(3): 583-596.

[23]
DANELLJAN M, BHAT G, KHAN F S, et al. Eco: effificient convolution operators for tracking[C]// CVPR. Proceedings of the Conference on Computer Vision and Pattern Recognition. New York: CVPR, 2017, 6931-6939.

[24]
LI B, WU W, WANG Q, et al. SiamRPN++: evolution of siamese visual tracking with very deep networks[C]// CVPR. Proceedings of the Conference on Computer Vision and Pattern Recognition. New York: CVPR, 2019: 4282-4291.

[25]
ZHAO M J, OKADA K, INABA M. TrTr: visual tracking with Transformer[EB/OL]. https://doi.org/CoRR2021.abs/2105.03817.

文章导航

/

[an error occurred while processing this directive]