[an error occurred while processing this directive] [an error occurred while processing this directive] [an error occurred while processing this directive]
[an error occurred while processing this directive]

基于编码记忆网络的半监督视频目标分割方法

  • 尹亮 1 ,
  • 张钊 2 ,
  • 张宝鹏 2
展开
  • 1 国能新朔铁路有限责任公司机务分公司,北京 010300
  • 2 北京交通大学计算机与信息技术学院,北京 100044
张宝鹏(1976—),男,副教授,博士,研究方向:计算机视觉,机器学习。

尹亮(1983—),男,工程师,研究方向:铁路机动检修。

收稿日期: 2024-02-20

  网络出版日期: 2024-12-28

基金资助

国家自然科学基金项目(61972027)

Method of Semi-supervised Video Object Segmentation Based on Encoding Memory Network

  • YIN Liang 1 ,
  • ZHANG Zhao 2 ,
  • ZHANG Baopeng 2
Expand
  • 1 CHN Energy Xinshuo Railway Co., Ltd. Maintenance Branch office, Beijing 010300, China
  • 2 School of Computer and Information Technology, Beijing Jiaotong University, Beijing 100044, China

Received date: 2024-02-20

  Online published: 2024-12-28

摘要

视频目标分割是计算机视觉中的一项关键任务,在自动驾驶、视频编码等领域具有重要意义。针对视频目标分割任务,提出使用一种高效的编码记忆网络(EMNet)实现半监督视频目标分割任务。该方法包含自适应参考帧选取模块、双路径匹配模块、特征处理模块以及特征聚合模块。自适应参考帧选取模块综合考虑掩码置信度和相似度,选择包含丰富信息的参考帧。双路径匹配模块实现查询帧和参考帧之间的双向和双尺度匹配,提高目标特征匹配准确率。特征处理模块分别包含语义强化模块和特征细化模块,通过低通和高通滤波增强目标的语义和细节信息。并由特征聚合模块对各特征进行融合利用。最后通过在DAVIS2017数据集上的评估,证明所提出模型的有效性。

本文引用格式

尹亮 , 张钊 , 张宝鹏 . 基于编码记忆网络的半监督视频目标分割方法[J]. 弹箭与制导学报, 2024 , 44(3) : 11 -21 . DOI: 10.15892/j.cnki.djzdxb.2024.03.002

Abstract

Video object segmentation is a key task in computer vision and is of great significance to fields such as autonomous driving and video coding. For the video object segmentation, the proposed method utilizes an efficient encoding memory network (EMNet) to achieve semi-supervised video object segmentation. The method includes an adaptive reference frame selection module, a dual path matching module, a feature processing module and a feature aggregation module. The adaptive reference frame selection module takes into account mask confidence and similarity, and selects a reference frame that contains rich information. The dual-path matching module realizes bidirectional and dual-scale matching between query frames and reference frames to improve the accuracy of target feature matching. The feature processing module includes a semantic enhancement module and a feature refinement module, which enhance the semantic and detailed information of the target through low-pass and high-pass filtering. Finally, the feature aggregation module fuses and utilizes each feature. An evaluation is carried out on the DAVIS2017dataset and the result shows that the proposed method is effective.

[an error occurred while processing this directive]
[1]
SIAM M, MAHGOUB H, ZAHRAN M, et al. Video object segmentation using speaker identification[C]// IEEE. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2018: 7386-7395.

[2]
LEE Y J, KIM J, GRAUMAN K. Key-segments for video object segmentation[C]// IEEE. Proceedings of the IEEE International Conference on Computer Vision. New York: IEEE, 2011: 1995-2002.

[3]
LI Y, YE Z, REHG J M. Delving into egocentric actions[C]// IEEE. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2015: 287-295.

[4]
HE K, GKIOXARI G, DOLLÁR P, et al. Mask r-cnn[C]// IEEE. Proceedings of the IEEE International Conference on Computer Vision. New York: IEEE, 2017: 2961-2969.

[5]
PERAZZI F, KHOREVA A, BENENSON R, et al. Learning video object segmentation from static images[C]// IEEE. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2017: 2663-2672.

[6]
TSAI Y H, YANG M H, BLACK M J. Video segmentation via object flow[C]// IEEE. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2016: 3899-3908.

[7]
OH S W, LEE J Y, XU N, et al. Video object segmentation using space-time memory networks[C]// IEEE. Proceedings of the IEEE/CVF International Conference on Computer Vision. New York: IEEE, 2019: 9226-9235.

[8]
JAIN M, JÉGOU H, BOUTHEMY P. Better exploiting motion for better action recognition[C]// IEEE. Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2013: 2555-2562.

[9]
MCFARLANE N J, SCHOFIELD C P. Segmentation and tracking of piglets in images[J]. Machine Vision and Applications, 1995, 8(3): 187-193.

[10]
GRUNDMANN M, KWATRA V, HAN M, et al. Efficient hierarchical graph-based video segmentation[C]// IEEE. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2010: 2141-2148.

[11]
OCHS P, BROX T. Higher order motion models and spectral clustering[C]// IEEE. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2012: 614-621.

[12]
ZHANG Z, HUANG K, TAN T. Comparison of similarity measures for trajectory clustering in outdoor surveillance scenes[C]// IEEE. Proceedings of the 18th International Conference on Pattern Recognition. New York: IEEE, 2006: 1135-1138.

[13]
HU W, HU M, ZHOU X, et al. Principal axis-based correspondence between multiple cameras for people tracking[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2006, 28(4): 663-671.

PMID

[14]
JAMPANI V, GADDE2 R, GEHLER P V. Video propagation networks[C]// IEEE. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2017: 451-461.

[15]
YANG L, WANG Y, XIONG X, et al. Efficient video object segmentation via network modulation[C]// IEEE. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2018: 6499-6507.

[16]
VOIGTLAENDER P, CHAI Y, SCHROFF F, et al. Feelvos: fast end-to-end embedding learning for video object segmentation[C]// IEEE. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2019: 9481-9490.

[17]
CAELLES S, MANINIS K K, PONT-TUSET J, et al. One-shot video object segmentation[C]// IEEE. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2017: 221-230.

[18]
LUITEN J, VOIGTLAENDER P, LEIBE B. Premvos: proposal-generation, refinement and merging for video object segmentation[C]// ACCV. Proceedings of the 2018 Asian Conference on Computer Vision. Berlin:Springer International Publishing, 2018: 565-580.

[19]
OH S W, LEE J Y, SUNKAVALLI K, et al. Fast video object segmentation by reference-guided mask propagation[C]// IEEE. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2018: 7376-7385.

[20]
WANG Z, XU J, LIU L, et al. Ranet: ranking attention network for fast video object segmentation[C]// IEEE. Proceedings of the IEEE/CVF International Conference on Computer Vision. New York: IEEE, 2019: 3978-3987.

[21]
WANG H, JIANG X, REN H, et al. Swiftnet: real-time video object segmentation[C]// IEEE. Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2021: 1296-1305.

[22]
SEONG H, HYUN J, KIM E. Kernelized memory network for video object segmentation[C]// ECCV. Proceedings of the 2020 Computer Vision-ECCV: 16th European Conference.Berlin:Springer International Publishing, 2020: 629-645.

[23]
IPPOLITI L, VALENTINI P, GAMERMAN D. Space-time modelling of coupled spatiotemporal environmental variables[J]. Journal of the Royal Statistical Society Series C: Applied Statistics, 2012, 61(2): 175-200.

[24]
LIN Z, YANG T, LI M, et al. Swem: Towards real-time video object segmentation with sequential weighted expectation-maximization[C]// IEEE. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2022: 1362-1372.

[25]
HU Y T, HUANG J B, SCHWING A G. Videomatch: Matching based video object segmentation[C]// ECCV. Proceedings of the 2018 European conference on computer vision (ECCV). Berlin:Springer International Publishing, 2018: 54-70.

[26]
HU L, ZHANG P, ZHANG B, et al. Learning position and target consistency for memory-based video object segmentation[C]// IEEE. Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2021: 4144-4154.

[27]
XU X, WANG J, LI X, et al. Reliable propagation-correction modulation for video object segmentation[C]// AIAA. Proceedings of the AIAA Conference on Artificial Intelligence. Reston: AIAA, 2022: 2946-2954.

[28]
ZHANG Y, ROBINSON A, MAGNUSSON M, et al. Flow-guided semi-supervised video object segmentation[C]// CVPR. Proceedings of the 2023 IEEE Computer Vision and Pattern Recognition Conference. New York: IEEE, 2023: 10492.

[29]
SEONG H, HYUN J, KIM E. Kernelized memory network for video object segmentation[C]// ECCV. Proceedings of the 2020 Computer Vision-ECCV 16th European Conference. Berlin:Springer International Publishing, 2020: 629-645.

文章导航

/

[an error occurred while processing this directive]