1. 西北机电工程研究所,陕西,咸阳,712099
2. 西安交通大学航天航空学院,西安,710049
3. 西安交通大学复杂服役环境重大装备结构强度与寿命全国重点实验室,西安,710049
: 2023-08-16。作者简介: 王鼎衡(1988—),男,助理研究员
杨朝旭(通信作者),男,副教授,硕士生导师。基金项目: 国家自然科学基金资助项目(12002254)。
网络首发:2024-03-10,
纸质出版:2024
移动端阅览
王鼎衡, 刘保荣, 杨维, 等. KCPStack:张量分解的卷积核分层矩阵压缩方法[J]. 西安交通大学学报, 2024,58(3):137-148.
WANG Dingheng, LIU Baorong, YANG Wei, et al. KCPStack: Tensor Decomposed Compression Method with Layered Matrices for Convolutional Kernels[J]. 2024, 58(3): 137-148.
王鼎衡, 刘保荣, 杨维, 等. KCPStack:张量分解的卷积核分层矩阵压缩方法[J]. 西安交通大学学报, 2024,58(3):137-148. DOI: 10.7652/xjtuxb202403013.
WANG Dingheng, LIU Baorong, YANG Wei, et al. KCPStack: Tensor Decomposed Compression Method with Layered Matrices for Convolutional Kernels[J]. 2024, 58(3): 137-148. DOI: 10.7652/xjtuxb202403013.
针对现有张量分解卷积核压缩方法难以兼顾时空轻量化、过于依赖卷积瓶颈结构等问题
提出一种具有可观压缩与加速能力的卷积核分层矩阵压缩方法(KCPStack)。首先
在矩阵乘法视角下
将卷积核按通道拆分为2阶克罗内克规范多项式(KCP)分解
所得因子张量组合为两层权重矩阵
使卷积计算转换为具有较高推理效率的双层轻量卷积结构; 其次
对比所提KCPStack方法与其他典型张量分解卷积核压缩方法的参数约减空间复杂度与推理计算时间复杂度; 最后
基于RK3588神经处理单元进行KCPStack方法的部署
面向实际场景目标检测识别需求开发相关应用。实验结果表明:与现有张量分解方法相比
在张量秩相同或者参数量相当的前提下
所提KCPStack方法具有最快的推理计算效率; 在图像分类标准数据集CIFAR-10和ImageNet上
KCPStack方法能够将精度损失控制在1%左右
最高可减少85.0%的参数量和79.8%的计算量; 在目标检测识别标准数据集COCO上
KCPStack方法相对于基线模型的平均精度下降不超过1%; 采用所提KCPStack方法对实际场景进行目标检测识别
在RK3588神经处理单元上能达到95.4%的平均精度和35帧/s的图像处理帧率
内存开销仅为33.1 MB。
To address the limitations of existing tensor decomposition-based methods for compressing convolutional kernels
such as the trade-off between spatial and temporal lightweight and excessive reliance on convolutional bottleneck structure
a compression method with layered matrices called KCPStack is proposed in this paper which offers considerable compression and acceleration capabilities. Firstly
from the perspective of matrix multiplication
the convolutional kernels are split by channel and subjected to a second-order Khatri-Rao Product(KCP)decomposition and the resulting factor tensors are combined into two-layer weight matrices
thereby transforming the convolutional computation into a two-layer lightweight convolutional structure with higher inference efficiency. Secondly
a comparison is made between the space complexity regarding parameter reduction and time complexity regarding inference computation of the KCPStack method and other typical tensor decomposition-based compression methods for convolutional kernels. Lastly
the KCPStack method is deployed on the RK3588 neural processing unit to develop related applications to meet the object detection and recognition needs for the real scene. The experimental results demonstrate that
compared with existing tensor decomposition-based methods
the proposed KCPStack method achieves the highest inference computation efficiency under the same tensor rank or comparable parameter quantity conditions. On the benchmark datasets CIFAR-10 and ImageNet for image classification
the KCPStack method controls the accuracy loss to around 1% while achieving a maximum parameter reduction of 85.0% and a computational saving of 79.8%. On the benchmark dataset COCO for object detection and recognition
the KCPStack method exhibits an average precision drop of less than 1% compared to the baseline model. When the KCPStack method is adopted for an object detection and recognition task in the real scene
an average precision of 95.4% and of an image processing frame rate of 35 frames per second are achieved on the RK3588 neural processing unit
and it requires only 33.1 MB of memory consumption.
DENG Lei, LI Guoqi, HAN Song, et al. Model compression and hardware acceleration for neural networks: a comprehensive survey [J]. Proceedings of the IEEE, 2020, 108(4): 485-532.
WU Yang, WANG Dingheng, LU Xiaotong, et al. Efficient visual recognition: a survey on recent advances and brain-inspired methodologies [J]. Machine Intelligence Research, 2022, 19(5): 366-411.
郭朝鹏, 王馨昕, 仲昭晋, 等. 能耗优化的神经网络轻量化方法研究进展 [J]. 计算机学报, 2023, 46(1): 85-102.
GUO Chaopeng, WANG Xinxin, ZHONG Zhaojin, et al. Research advance on neural network lightweight for energy optimization [J]. Chinese Journal of Computers, 2023, 46(1): 85-102.
林景栋, 吴欣怡, 柴毅, 等. 卷积神经网络结构优化综述 [J]. 自动化学报, 2020, 46(1): 24-37.
LIN Jingdong, WU Xinyi, CHAI Yi, et al. Structure optimization of convolutional neural networks: a survey [J]. Acta Automatica Sinica, 2020, 46(1): 24-37.
WANG Maolin, PAN Yu, XU Zenglin, et al. Tensor networks meet neural networks: a survey and future perspectives [EB/OL].(2023-05-08)[2023-07-01]. https://arxiv.org/abs/2302.09019.
WANG Dingheng, ZHAO Guangshe, CHEN Hengnu, et al. Nonlinear tensor train format for deep neural network compression [J]. Neural Networks, 2021, 144: 320-333.
NOVIKOV A, PODOPRIKHIN D, OSOKIN A, et al. Tensorizing neural networks [C]//Proceedings of the 28th International Conference on Neural Information Processing Systems. Cambridge, MA, USA: MIT Press, 2015: 442-450.
KOSSAIFI J, TOISOUL A, BULAT A, et al. Factorized higher-order CNNs with an application to spatio-temporal emotion estimation [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2020: 6059-6068.
YANG Yinchong, KROMPASS D, TRESP V. Tensor-train recurrent neural networks for video classification [C]//Proceedings of the 34th International Conference on Machine Learning. Piscataway, NJ, USA: IEEE, 2017: 3891-3900.
YE Jinmian, WANG Linnan, LI Guangxi, et al. Learning compact recurrent neural networks with block-term tensor decomposition [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 9378-9387.
PAN Yu, XU Jing, WANG Maolin, et al. Compressing recurrent neural networks with tensor ring for action recognition [C]//Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto, CA, USA: AAAI Press, 2019: 4683-4690.
YIN Miao, LIAO Siyu, LIU Xiaoyang, et al. Towards extremely compact RNNs for video recognition with fully decomposed hierarchical tucker structure [C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2021: 12080-12089.
李大鹏, 陈剑, 王晨, 等. 基于模板张量分解和双向LS TM的司法案件罪名认定 [J]. 电子学报, 2021, 49(4): 760-767.
LI Dapeng, CHEN Jian, WANG Chen, et al. Conviction in judicial cases based on template tensor decomposition and bidirectional LSTM [J]. Acta Electronica Sinica, 2021, 49(4): 760-767.
李晶晶, 夏鸿斌, 刘渊. 融合注意力LSTM的神经张量分解推荐模型 [J]. 中文信息学报, 2021, 35(5): 91-100.
LI Jingjing, XIA Hongbin, LIU Yuan. Neural tensor factorization recommendation model based on attention LSTM [J]. Journal of Chinese Information Processing, 2021, 35(5): 91-100.
ZHANG Xiangyu, ZOU Jianhua, MING Xiang, et al. Efficient and accurate approximations of nonlinear convolutional networks [C]//2015 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2015: 1984-1992.
KIM Y D, PARK E, YOO S, et al. Compression of deep convolutional neural networks for fast and low power mobile applications [C]//International Conference on Learning Representations(ICLR). San Juan, Puerto Rico: Open Review, 2016: 1-16.
CHEN Yunpeng, JIN Xiaojie, KANG Bingyi, et al. Sharing residual units through collective tensor factorization to improve deep neural networks [C]//Proceedings of the 27th International Joint Conference on Artificial Intelligence. Palo Alto, CA, USA: AAAI Press, 2018: 635-641.
ASTRID M, LEE S I. CP-decomposition with tensor power method for convolutional neural networks compression [C]//2017 IEEE International Conference on Big Data and Smart Computing(BigComp). Piscataway, NJ, USA: IEEE, 2017: 115-118.
LEBEDEV V, GANIN Y, RAKHUBA M, et al. Speeding-up convolutional neural networks using fine-tuned CP-decomposition [C]//International Conference on Learning Representations(ICLR). San Diego, CA, USA: Open Review, 2015: 1-11.
王鼎衡, 赵广社, 姚满, 等. KCPNet:张量分解的轻量卷积模块设计、部署与应用 [J]. 西安交通大学学报, 2022, 56(3): 135-146.
WANG Dingheng, ZHAO Guangshe, YAO Man, et al. KCPNet: design, deployment, and application of tensor-decomposed lightweight convolutional module [J]. Journal of Xi'an Jiaotong University, 2022, 56(3): 135-146.
WANG Dingheng, WU Bijiao, ZHAO Guangshe, et al. Kronecker CP decomposition with fast multiplication for compressing RNNs [J]. IEEE Transactions on Neural Networks and Learning Systems, 2023, 34(5): 2205-2219.
PHAN A H, CICHOCKI A, TICHAVSK P, et al. From basis components to complex structural patterns [C]//2013 IEEE International Conference on Acoustics, Speech and Signal Processing. Piscataway, NJ, USA: IEEE, 2013: 3228-3232.
CHETLUR S, WOOLLEY C, VANDERMERSCH P, et al. cuDNN: efficient primitives for deep learning [EB/OL].(2014-12-18)[2023-07-01]. https://arxiv.org/abs/1410.0759.
WU Bijiao, WANG Dingheng, ZHAO Guangshe, et al. Hybrid tensor decomposition in neural network compression [J]. Neural Networks, 2020, 132: 309-320.
DOLGOV S V, SAVOSTYANOV D V. Alternating minimal energy methods for linear systems in higher dimensions [J]. SIAM Journal on Scientific Computing, 2014, 36(5): A2248-A2271.
SANDLER M, HOWARD A, ZHU Menglong, et al. MobileNetV2: inverted residuals and linear bottlenecks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 4510-4520.
0
浏览量
7
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621