

浏览全部资源
扫码关注微信
西安交通大学自动化科学与工程学院,西安,710049
Online First:10 March 2022,
Published:2022
移动端阅览
MRTP:Multi-Temporal Resolution Real-Time Action Recognition Approach by Time-Action Perception[J]. 2022, 56(3): 22-32.
MRTP:Multi-Temporal Resolution Real-Time Action Recognition Approach by Time-Action Perception[J]. 2022, 56(3): 22-32. DOI: 10.7652/xjtuxb202203003.
针对行为识别中时空信息分布不均衡以及对长时间跨度信息表征获取难的问题
提出了一种时间-动作感知的多尺度时间序列实时行为识别方法MRTP。以RGB视频为输入
使用两个并行的感知路径在不同的时间分辨率上对视频进行空间特征与动作特征提取。在空间路径中
使用基于特征差分的动作感知寻找并加强通道动作特征表征; 在动作路径中
基于动作感知的权重对通道进行筛选
并加入通道注意力和时间注意力加强关键特征; 在两个路径提取出特征后
对特征进行融合
融合后的特征通过激活函数映射出样本在各个类别的得分
取得分最高的类别为最终识别结果。实验结果表明:所提方法在UCF101数据集上达到了95.6%的准确率
优于未使用时间注意力的方法; 在AVA2.2数据集上的平均精度达到了28%
优于未使用动作感知和时间注意力的方法。与目前主流的基于光流法的双流网络、以Slowfast为代表的3D卷积网络、Transformer等方法进行了准确率、参数量、处理速度对比
结果表明所提方法具有更良好的识别效果和鲁棒性。
In view of the uneven distribution of temporal and spatial information and the difficulty in obtaining long-term information representation
we propose a dual-path action recognition method MRTP based on time and motion perception
which takes RGB video as input and applies two parallel perception paths to extract spatial and motion features from the video at different time resolutions. In the spatial path
the action-perception based on feature difference is applied to find and strengthen the action feature representation; in the action path
the channel is filtered based on the weight of action perception
and channel attention and time attention are added to enhance the key features; the characteristics of the two paths are merged to calculate the action category score of the video. The experimental results show that MRTP achieves accuracy of 95.6% on the UCF101 data set to outperform the model without time attention; on the AVA2.2 data set
the accuracy of mAP reaches 28% to outperform the model without action perception and time attention. Compared with the current mainstream two-stream network
3D convolution
Transformer
and other methods on a number of accuracy indicators
this method is endowed with better recognition effect and robustness.
朱煜, 赵江坤, 王逸宁, 等. 基于深度学习的人体行为识别算法综述 [J]. 自动化学报, 2016, 42(6): 848-857.
ZHU Yu, ZHAO Jiangkun, WANG Yining, et al. A review of human action recognition based on deep learning [J]. Acta Automatica Sinica, 2016, 42(6): 848-857.
罗会兰, 王婵娟, 卢飞. 视频行为识别综述 [J]. 通信学报, 2018, 39(6): 169-180.
LUO Huilan, WANG Chanjuan, LU Fei. Survey of video behavior recognition [J]. Journal on Communications, 2018, 39(6): 169-180.
ZHU Yi, LI Xinyu, LIU Chunhui, et al. A comprehensive study of deep video action recognition [EB/OL]. [2021-04-01]. https: ∥arxiv.org/abs/2012. 06567.
VAROL G, LAPTEV I, SCHMID C. Long-term temporal convolutions for action recognition [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(6): 1510-1517.
CHEN Yuehai, YANG Jing, ZHANG Kun, et al. A feature-cascaded correntropy LSTM for tourists prediction [J]. IEEE Access, 2021, 9: 32810-32822.
LI Zhenyang, GAVRILYUK K, GAVVES E, et al. VideoLSTM convolves, attends and flows for action recognition [J]. Computer Vision and Image Understanding, 2018, 166: 41-50.
CARREIRA J, ZISSERMAN A. Quo vadis, action recognition? a new model and the kinetics dataset [C]∥Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2017: 4724-4733.
FEICHTENHOFER C, FAN Haoqi, MALIK J, et al. SlowFast networks for video recognition [C]∥Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2019: 6201-6210.
ZHANG Shiwen, GUO Sheng, HUANG Weilin, et al. V4D: 4D convolutional neural networks for video-level representation learning [C/OL]∥Proceedings of the 2020 International Conference on Learning Representations. London, UK: ICLR, 2020 [2021-04-01]. https: ∥arxiv.org/pdf/2002.07442.pdf.
LAPTEV I. On space-time interest points [J]. International Journal of Computer Vision, 2005, 64(2): 107-123.
DOLLAR P, RABAUD V, COTTRELL G, et al. Behavior recognition via sparse spatio-temporal features [C]∥Proceedings of the 2005 IEEE International Workshop on Visual Surveillance and Performance Evaluation of Tracking and Surveillance. Piscataway, NJ, USA: IEEE, 2005: 65-72.
BOBICK A F, DAVIS J W. The recognition of human movement using temporal templates [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2001, 23(3): 257-267.
LAPTEV I, MARSZALEK M, SCHMID C, et al. Learning realistic human actions from movies [C]∥Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2008: 1-8.
LAZEBNIK S, SCHMID C, PONCE J. Beyond bags of features: spatial pyramid matching for recognizing natural scene categories [C]∥Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2006: 2169-2178.
PERRONNIN F, DANCE C. Fisher kernels on visual vocabularies for image categorization [C]∥Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2007: 1-8.
WANG Heng, SCHMID C. Action recognition with improved trajectories [C]∥Proceedings of the 2013 IEEE International Conference on Computer Vision. Piscataway, NJ, USA: IEEE, 2013: 3551-3558.
WANG Limin, XIONG Yuanjun, WANG Zhe, et al. Temporal segment networks: towards good practices for deep action recognition [C]∥Proceedings of the 2016 European Conference on Computer Vision. Cham, Germany: Springer, 2016: 20-36.
FEICHTENHOFER C, PINZ A, ZISSERMAN A. Convolutional two-stream network fusion for video action recognition [C]∥Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2016: 1933-1941.
ILG E, MAYER N, SAIKIA T, et al. FlowNet 2.0: evolution of optical flow estimation with deep networks [C]∥Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2017: 1647-1655.
SIMONYAN K, ZISSERMAN A. Two-stream convolutional networks for action recognition in videos [C]∥Proceedings of the 27th International Conference on Neural Information Processing Systems. Vancouver, Canada: NIPS, 2014: 568-576.
HOCHREITER S, SCHMIDHUBER J. Long short-term memory [J]. Neural Computation, 1997, 9(8): 1735-1780.
NG J Y H, HAUSKNECHT M, VIJAYANARASIMHAN S, et al. Beyond short snippets: deep networks for video classification [C]∥Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2015: 4694-4702.
MA C Y, CHEN M H, KIRA Z, et al. TS-LSTM and temporal-inception: exploiting spatiotemporal dynamics for activity recognition [J]. Signal Processing: Image Communication, 2019, 71: 76-87.
TRAN D, BOURDEV L, FERGUS R, et al. Learning spatiotemporal features with 3D convolutional networks [C]∥Proceedings of the 2015 IEEE International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2015: 4489-4497.
KAY W, CARREIRA J, SIMONYAN K, et al. The kinetics human action video dataset [EB/OL]. [2021-04-01]. https: ∥arxiv.org/abs/1705.06950.
QIU Zhaofan, YAO Ting, MEI Tao. Learning spatio-temporal representation with pseudo-3D residual networks [C]∥Proceedings of the 2017 IEEE International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2017: 5534-5542.
HARA K, KATAOKA H, SATOH Y. Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet? [C]∥Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 6546-6555.
SOOMRO K, ZAMIR A R, SHAH M. UCF101: a dataset of 101 human actions classes from videos in the wild [EB/OL]. [2021-04-01]. https: ∥arxiv.org/abs /1212.0402.
GU Chunhui, SUN Chen, ROSS D A, et al. AVA: a video dataset of spatio-temporally localized atomic visual actions [C]∥Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 6047-6056.
FEICHTENHOFER C. X3D: expanding architectures for efficient video recognition [C]∥Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2020: 200-210.
WANG Xianyuan, MIAO Zhenjiang, ZHANG Ruyi, et al. I3D-LSTM: a new model for human action recognition [C]∥Proceedings of the IOP Conference Series: Materials Science and Engineering. London, UK: IOP, 2019: 032035.
HUANG Guoxi, BORS A G. Learning spatio-temporal representations with temporal squeeze pooling [C]∥Proceedings of the 2020 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP). Piscataway, NJ, USA: IEEE, 2020: 2103-2107.
SUN Chen, SHRIVASTAVA A, VONDRICK C, et al. Actor-centric relation network [C]∥Proceedings of the 2018 European Conference on Computer Vision. Cham, Germany: Springer, 2018: 335-351.
GIRDHAR R, CARREIRA J, DOERSCH C, et al. A better baseline for AVA [EB/OL]. [2021-04-01]. https: ∥arxiv.org/abs/1807.10066.
GIRDHAR R, JO(~overA)O CARREIRA J, DOERSCH C, et al. Video action transformer network [C]∥Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2019: 244-253.
STROUD J C, ROSS D A, SUN Chen, et al. D3D: distilled 3D networks for video action recognition [C]∥Proceedings of the 2020 IEEE Winter Conference on Applications of Computer Vision(WACV). Piscataway, NJ, USA: IEEE, 2020: 614-623.
WANG Xiaolong, GIRSHICK R, GUPTA A, et al. Non-local neural networks [C]∥Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 7794-7803.
TRAN D, WANG Heng, TORRESANI L, et al. A closer look at spatiotemporal convolutions for action recognition [C]∥Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 6450-6459.
0
Views
7
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024621