

浏览全部资源
扫码关注微信
西安交通大学自动化科学与工程学院,西安,710049
Online First:10 March 2022,
Published:2022
移动端阅览
KCPNet:Design,Deployment,and Application of Tensor-Decomposed Lightweight Convolutional Module[J]. 2022, 56(3): 135-146.
KCPNet:Design,Deployment,and Application of Tensor-Decomposed Lightweight Convolutional Module[J]. 2022, 56(3): 135-146. DOI: 10.7652/xjtuxb202203014.
为解决现有卷积模块在实际应用中内存消耗高、计算效率低的问题
在Kronecker CANDECOMP/PARAFAC(KCP)张量分解的基础上
提出一种轻量、高效、瓶颈结构的卷积模块(KCPNet)。对普通卷积作2阶KCP分解
生成的因子张量分别映射为两层负责输入输出通道变化的1×1卷积和两层负责特征提取的变通道可分离卷积
再将这4层卷积组成含有瓶颈结构的KCPNet卷积模块。基于OpenCL并行编程框架将KCPNet部署于嵌入式GPU
并围绕pico-flexx深度相机开发了动态手势识别应用。实验结果表明:在ImageNet大规模标准数据集上
相比ResNet、ResNeXt等已有的张量分解卷积模块
KCPNet在准确率相近的情况下能够兼顾空间和计算复杂度的效率; 在中等规模标准数据集CIFAR-10上
KCPNet能够在无明显精度损失的前提下将传统的VGG模型压缩至原先的16.1%并节约75.5%的计算量; 在面向嵌入式GPU时
并行部署的KCPNet可使CIFAR-10的识别速度达到100帧/s。以KCPNet为核心开发的手势识别应用程序可达到99.5%的准确率和100帧/s以上的运行速度
内存开销为22 MB。
To deal with the issues of high memory consumption and low computation efficiency of the existing convolutional modules in practical application
following the Kronecker CANDECOMP/PARAFAC(KCP)tensor decomposition
a lightweight
efficient
bottleneck-structured convolutional module(KCPNet)is proposed. The normal convolution is firstly decomposed as the 2nd-order KCP
the generated factor tensors are mapped into two 1×1 convolutions to carry out the transformation of input and output channels
and into two depthwise convolutions with changeable channels to perform the feature extraction
respectively. Then
these 4 convolutions are assembled into the KCPNet module with a bottleneck structure. The KCPNet is deployed in the embedded GPU based on the OpenCL parallel programming framework
and the dynamic gesture recognition application is developed around a pico-flexx depth camera. The experimental results show that compared with the existing convolutional modules of tensor decomposition such as ResNet
ResNeXt
etc.
the KCPNet can consider the efficiency of both space and computation complexity with a similar accuracy on the large-scale benchmark dataset ImageNet; the KCPNet can also compress the traditional VGG model to its original 16.1% and save 75.5% of computation without obvious accuracy loss on the medium-scale benchmark dataset CIFAR-10; the parallelly deployed KCPNet can achieve the recognition rate of 100 frames/s for CIFAR-10 when orienting to the embedded GPU. The developed gesture recognition application with the KCPNet as its core can achieve 99.5% accuracy and the running rate beyond 100 frames/s
and the memory cost is 22 MB.
LECUN Y, BENGIO Y, HINTON G. Deep learning [J]. Nature, 2015, 521(7553): 436-444.
DENG Lei, LI Guoqi, HAN Song, et al. Model compression and hardware acceleration for neural networks: a comprehensive survey [J]. Proceedings of the IEEE, 2020, 108(4): 485-532.
林景栋, 吴欣怡, 柴毅, 等. 卷积神经网络结构优化综述 [J]. 自动化学报, 2020, 46(1): 24-37.
LIN Jingdong, WU Xinyi, CHAI Yi, et al. Structure optimization of convolutional neural networks: a survey [J]. Acta Automatica Sinica, 2020, 46(1): 24-37.
毛远宏, 贺占庄, 刘露露. 目标跟踪中基于深度可分离卷积的剪枝方法 [J]. 西安交通大学学报, 2021, 55(1): 52-59.
MAO Yuanhong, HE Zhanzhuang, LIU Lulu. Pruning based on separable convolutions for object tracking [J]. Journal of Xi'an Jiaotong University, 2021, 55(1): 52-59.
LIN Shaohui, JI Rongrong, LI Yuchao, et al. Toward compact ConvNets via structure-sparsity regularized filter pruning [J]. IEEE Transactions on Neural Networks and Learning Systems, 2020, 31(2): 574-588.
梁峰, 董名, 田志超, 等. 面向轻量化神经网络的模型压缩与结构搜索 [J]. 西安交通大学学报, 2020, 54(11): 106-112.
LIANG Feng, DONG Ming, TIAN Zhichao, et al. Model compression and structure search for lightweight neural network [J]. Journal of Xi'an Jiaotong University, 2020, 54(11): 106-112.
WANG Tianzhe, WANG Kuan, CAI Han, et al. APQ: joint search for network architecture, pruning and quantization policy [C]∥Proceeding of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2020: 2075-2084.
程士卿, 郝问裕, 李晨, 等. 低秩张量分解的多视角谱聚类算法 [J]. 西安交通大学学报, 2020, 54(3): 119-125, 133.
CHENG Shiqing, HAO Wenyu, LI Chen, et al. Multi-view clustering by low-rank tensor decomposition [J]. Journal of Xi'an Jiaotong University, 2020, 54(3): 119-125, 133.
AGGARWAL V, WANG Wenlin, ERIKSSON B, et al. Wide compression: tensor ring nets [C]∥Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 9329-9338.
WU Bijiao, WANG Dingheng, ZHAO Guangshe, et al. Hybrid tensor decomposition in neural network compression [J]. Neural Networks, 2020, 132: 309-320.
HUANG Hantao, NI Leibin, WANG Kanwen, et al. A highly parallel and energy efficient three-dimensional multilayer CMOS-RRAM accelerator for tensorized neural network [J]. IEEE Transactions on Nanotechnology, 2018, 17(4): 645-656.
DENG Chunhua, SUN Fangxuan, QIAN Xuehai, et al. TIE: energy-efficient tensor train-based inference engine for deep neural network [C]∥Proceedings of the 46th International Symposium on Computer Architecture. New York, USA: ACM, 2019: 264-278.
KOSSAIFI J, TOISOUL A, BULAT A, et al. Factorized higher-order CNNs with an application to spatio-temporal emotion estimation [C]∥Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2020: 6059-6068.
ZHOU Shuchang, WU Jianan, WU Yuxin, et al. Exploiting local structures with the Kronecker layer in convolutional networks [EB/OL].(2016-02-04)[2021-07-01]. https: ∥arxiv.org/abs/1512.09194.
HAMEED MGA, TAHAEI MS, MOSLEH A, et al. Convolutional neural network compression through generalized Kronecker product decomposition [EB/OL].(2021-09-29)[2021-10-13]. https: ∥arxiv.org/abs/2109.14710.
WANG Dingheng, ZHAO Guangshe, LI Guoqi, et al. Compressing 3DCNNs based on tensor train decomposition [J]. Neural Networks, 2020, 131: 215-230.
LEE D, WANG Dingheng, YANG Yukuan, et al. QTTNet: quantized tensor train neural networks for 3D object and video recognition [J]. Neural Networks, 2021, 141: 420-432.
ZHANG Xiangyu, ZOU Jianhua, MING Xiang, et al. Efficient and accurate approximations of nonlinear convolutional networks [C]∥Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2015: 1984-1992.
LEBEDEV V, GANIN Y, RAKHUBA M, et al. Speeding-up convolutional neural networks using fine-tuned CP-decomposition [C]∥Proceedings of the 3rd International Conference on Learning Representations. London, UK: ICLR, 2015: 1-11.
ASTRID M, LEE S I. CP-decomposition with tensor power method for convolutional neural networks compression [C]∥Proceedings of the 2017 IEEE International Conference on Big Data and Smart Computing(BigComp). Piscataway, NJ, USA: IEEE, 2017: 115-118.
KIM Y D, PARK E, YOO S, et al. Compression of deep convolutional neural networks for fast and low power mobile applications [C]∥Proceedings of the 4th International Conference on Learning Representations. London, UK: ICLR, 2016: 1-16.
HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition [C]∥Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2016: 770-778.
CHEN Yunpeng, JIN Xiaojie, KANG Bingyi, et al. Sharing residual units through collective tensor factorization to improve deep neural networks [C]∥Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. Sacramento, CA, USA: IJCAI, 2018: 635-641.
XIE Saining, GIRSHICK R, DOLLÁR P, et al. Aggregated residual transformations for deep neural networks [C]∥Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2017: 5987-5995.
SANDLER M, HOWARD A, ZHU Menglong, et al. MobileNetV2: inverted residuals and linear bottlenecks [C]∥Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 4510-4520.
SUN Ke, LI Mingjie, LIU Dong, et al. IGCV3: interleaved low-rank group convolutions for efficient deep neural networks [C]∥Proceedings of the 2018 British Machine Vision Conference. Guildford, UK: BMVA Press, 2018: 1-13.
王子愉, 袁春, 黎健成. 利用可分离卷积和多级特征的实例分割 [J]. 软件学报, 2019, 30(4): 954-961.
WANG Ziyu, YUAN Chun, LI Jiancheng. Instance segmentation with separable convolutions and multi-level features [J]. Journal of Software, 2019, 30(4): 954-961.
周云成, 许童羽, 邓寒冰, 等. 基于面向通道分组卷积网络的番茄主要器官实时识别 [J]. 农业工程学报, 2018, 34(10): 153-162.
ZHOU Yuncheng, XU Tongyu, DENG Hanbing, et al. Real-time recognition of main organs in tomato based on channel wise group convolutional network [J]. Transactions of the Chinese Society of Agricultural Engineering, 2018, 34(10): 153-162.
杨贤志, 黄国方, 周宁宁. 基于分组卷积和特征图级联的轻量级目标检测 [J]. 计算机应用研究, 2021, 38(5): 1590-1594.
YANG Xianzhi, HUANG Guofang, ZHOU Ningning. Light-weight object detection network based on group convolution and feature maps cascade [J]. Application Research of Computers, 2021, 38(5): 1590-1594.
PHAN A H, CICHOCKI A, TICHAVSKÝ P, et al. From basis components to complex structural patterns [C]∥Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing. Piscataway, NJ, USA: IEEE, 2013: 3228-3232.
LEE N, CICHOCKI A. Regularized computation of approximate pseudoinverse of large matrices using low-rank tensor train decompositions [J]. SIAM Journal on Matrix Analysis and Applications, 2016, 37(2): 598-623.
RUSSAKOVSKY O, DENG Jia, SU Hao, et al. ImageNet large scale visual recognition challenge [J]. International Journal of Computer Vision, 2015, 115(3): 211-252.
ITO Y, MATSUMIYA R, ENDO T. ooc_cuDNN: accommodating convolutional neural networks over GPU memory capacity [C]∥Proceedings of the 2017 IEEE International Conference on Big Data. Piscataway, NJ, USA: IEEE, 2017: 183-192.
PRESNOV D, LAMBERS M, KOLB A. Robust range camera pose estimation for mobile online scene reconstruction [J]. IEEE Sensors Journal, 2018, 18(7): 2903-2915.
Melexis Inspired Engineering. Melexis to participate in the MinTOFKA project to develop a miniaturized 3D camera system for vehicle interiors [EB/OL].(2019-03-25)[2021-07-01]. https: ∥www. melexis.com/zh/news/2019/25 mar2019-mintofka-pro ject-kick-off.
PMD Technologies. Picofamily website relaunch and new online shop [EB/OL].(2018-03-13)[2021-07-01]. https: ∥pmdtec.com/en/company/news/picofamily-website-relaunch-and-new-online-shop/.
王鼎衡, 赵广社, 李国齐, 等. 基于张量链压缩的卷积神经网络及手势识别应用研究 [C]∥2018中国自动化大会论文集. 北京: 中国自动化学会, 2018: 87-92.
0
Views
5
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024621