In order to address the compatibility issue between model compression algorithms and the versatile tensor accelerator(VTA)
an adaptive fine-grained structured sparse design tailored for this accelerator is proposed by enhancing the classical YOLObile block-wise pruning method and evaluates its performance. In light of the multi-dimensional loop unfolding characteristics of VTA
the model's weight tensors are divided into 32×32 blocks. This approach integrates temporal distillation and spatial distillation to align multidimensional features. Through a single-stage iterative training method
the calculation process of the original ADMM algorithm is refined to improve model deployment accuracy while reducing training costs. An adaptive layer pruning rate module is introduced to dynamically allocate the total pruning rate
facilitating end-to-end automated pruning. The experimental results demonstrate that this improved method effectively reduces floating-point computations by approximately 2.4% and enhances the accuracy of compressed models across various tasks such as image classification and object detection
with a maximum growth percentage of 2.6%. This method offers an efficient and lightweight software solution for the sparse deployment of deep learning models on VTAs.
GAO Han, TIAN Yulong, XU Fengyuan, et al. Survey of deep learning model compression and acceleration [J]. Journal of Software, 2021, 32(1): 68-92.
BERTHELIER A, CHATEAU T, DUFFNER S, et al. Deep model compression and architecture optimization for embedded systems: a survey [J]. Journal of Signal Processing Systems, 2021, 93(8): 863-878.
FU Huitong, WANG Peng, LI Xiaoyan, et al. Lightweight network model for moving object recognition [J]. Journal of Xi'an Jiaotong University, 2021, 55(7): 124-131.
BA L J, CARUANA R. Do deep nets really need to be deep? [C]//Proceedings of the 27th International Conference on Neural Information Processing Systems. Cambridge, MA, USA: MIT Press, 2014: 2654-2662.
KIM T, KWON Y, LEE J, et al. CPrune: compiler-informed model pruning for efficient target-aware DNN execution [C]//Computer Vision -ECCV 2022. Cham, Switzerland: Springer Nature, 2022: 651-667.
LI Zhengang, YUAN Geng, NIU Wei, et al. NPAS: A compiler-aware framework of unified network pruning and architecture search for beyond real-time mobile acceleration [C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2021: 14250-14261.
GUAN Hui, LIU Shaoshan, MA Xiaolong, et al. CoCoPIE: enabling real-time AI on off-the-shelf mobile devices via compression-compilation co-design [J]. Communications of the ACM, 2021, 64(6): 62-68.
CHEN Tianqi, MOREAU T, JIANG Ziheng, et al. TVM: an automated end-to-end optimizing compiler for deep learning [C]//Proceedings of the 13th USENIX conference on Operating Systems Design and Implementation. USA: USENIX Association, 2018: 579-594.
MOREAU T, CHEN Tianqi, VEGA L, et al. A hardware-software blueprint for flexible deep learning specialization [J]. IEEE Micro, 2019, 39(5): 8-16.
Mudigere D, HAO Yuchen, HUANG Jianyu, et al. Software-hardware co-design for fast and scalable training of deep learning recommendation models [C]//Proceedings of the 49th Annual International Symposium on Computer Architecture. New York, USA: ACM, 2022: 993-1011.
VADERA S, AMEEN S. Methods for pruning deep neural networks [J]. IEEE Access, 2022, 10: 63280-63300.
HUANG Sitao, PEARSON C, NAGI R, et al. Accelerating sparse deep neural networks on FPGAs [C]//2019 IEEE High Performance Extreme Computing Conference(HPEC). Piscataway, NJ, USA: IEEE, 2019: 1-7.
ZHOU Aojun, MA Yukun, ZHU Junnan, et al. Learning N:M fine-grained structured sparse neural networks from scratch [EB/OL].(2021-04-18)[2024-04-01]. https://arxiv.org/abs/2102.04010.
CHANG S E, LI Yanyu, SUN Mengshu, et al. Mix and match: a novel FPGA-centric deep neural network quantization framework [C]//2021 IEEE International Symposium on High-Performance Computer Architecture(HPCA). Piscataway, NJ, USA: IEEE, 2021: 208-220.
LIN Jingdong, WU Xinyi, CHAI Yi, et al. Structure optimization of convolutional neural networks: a survey [J]. Acta Automatica Sinica, 2020, 46(1): 24-37.
CAI Yuxuan, LI Hongjia, YUAN Geng, et al. YOLObile: real-time object detection on mobile devices via compression-compilation co-design [C]//Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto, CA, USA: AAAI Press, 2021: 955-963.
FARHADI A, REDMON J. Yolov3: an incremental improvement [C]//Computer Vision and Pattern Recognition. Berlin/Heidelberg, Germany: Springer, 2018: 1-6.
HINTON G, VINYALS O, DEAN J. Distilling the knowledge in a neural network [EB/OL].(2015-03-09)[2024-04-01]. https://arxiv.org/abs/1503.02531.
NIU Wei, LI Zhengang, MA Xiaolong, et al. GRIM: a general, real-time deep learning inference framework for mobile devices based on fine-grained structured weight sparsity [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(10): 6224-6239.
BOCHKOVSKIY A, WANG C Y, LIAO H Y M. YOLOv4: optimal speed and accuracy of object detection [EB/OL].(2020-04-23)[2024-04-01]. https://arxiv.org/abs/2004.10934.
BAE W, YOO J, YE J C. Beyond deep residual learning for image restoration: persistent homology-guided manifold simplification [C]//2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops(CVPRW). Piscataway, NJ, USA: IEEE, 2017: 1141-1149.
SANDLER M, HOWARD A, ZHU Menglong, et al. MobileNetV2: inverted residuals and linear bottlenecks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 4510-4520.
HU Jie, SHEN Li, SUN Gang. Squeeze-and-excitation networks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 7132-7141.
HUANG Gao, LIU Zhuang, VAN DER MAATEN L, et al. Densely connected convolutional networks [C]//2017 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2017: 2261-2269.
FANG Gongfan, MA Xinyin, SONG Mingli, et al. DepGraph: towards any structural pruning [C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2023: 16091-16101.
SUN Yongshuai, GUO Mengyu, LIANG Dacheng, et al. Exploiting dynamic bit sparsity in activation for deep neuralnetwork acceleration [C]//2021 IEEE 14th International Conference on ASIC(ASICON). Piscataway, NJ, USA: IEEE, 2021: 1-4.