A Deep Learning Frame on Embedded Multicore Processors Based on Caffe and Its Parallel Implementation[J]. 2018, 52(6): 36-41+113.
DOI:
A Deep Learning Frame on Embedded Multicore Processors Based on Caffe and Its Parallel Implementation[J]. 2018, 52(6): 36-41+113.DOI: 10.7652/xjtuxb201806006.
A Deep Learning Frame on Embedded Multicore Processors Based on Caffe and Its Parallel Implementation
An effective embedded homogeneous and heterogeneous parallel improvement design on the basis of Caffe is proposed to solve the poor compatibility and low efficiency of forward inference in Android mobile terminals by using the open-sourced deep learning frame named Caffe(Convolutional architecture for fast feature embedding). The scheme transplants Caffe and its third-party library to arm architecture using a cross compiler
and then the multi-core and multi-thread technology is used to parallelize partial forward inference between convolution layer and input frame group. An heterogeneous parallel convolutional implementation based on OpenCL is also presented to further improve the time performance of the scheme. Comparison tests with three classic deep learning neural networks MNIST
Cifar-10 and CaffeNet show that in the absence of any model precision loss
the time consuming after parallelization is far less than that before parallel
and time performance increases up to 2 times. It is concluded that the proposal can make the deep learning frame Caffe effectively deploy and work in parallel on portable embedded multicore devices.
关键词
Keywords
references
JIA Yangqing, SHELHAMER E, DONAHUE J, et al. Caffe: convolutional architecture for fast feature embedding [C]∥Proceedings of the 22nd ACM International Conference on Multimedia. New York, USA: ACM, 2014: 675-678.
Google Inc. Tensorflow [EB/OL].(2017-09-23)[2017-10-10]. https: ∥github.com/tensorflow/tenso-rflow.
Facebook Inc. caffe2 [EB/OL].(2017-05-22)[2017-06-24]. https: ∥github.com/caffe2/caffe2.
Tencent Inc. ncnn [EB/OL].(2017-09-23)[2017-11-04]. https: ∥github.com/Tencent/ncnn.
Baidu Inc. Mobile-deep-learning [EB/OL].(2017-09-03)[2017-10-02]. https: ∥github.com/baidu/mobi-le-deep-learning.
HAN Song, MAO Huizi, DALLY W J. Deep compression: compressing deep neural networks with pruning, trained quantization and Huffman coding [J]. Fiber, 2015, 56(4): 3-7.
LIN D D, TALATHI S S, ANNAPUREDDY V S. Fixed point quantization of deep convolutional networks [J]. Computer Science, 2016: arXiv: 1511. 06393.
MARAT D. NNPACK [EB/OL].(2017-07-21)[2017-09-07]. https: ∥github.com/Maratyszcza/NNPACK.
OSKOUEI S S L, GOLESTANI H, KACHUEE M, et al. GPU-based acceleration of deep convolutional neural networks on mobile platforms [J]. Computer Science, 2015: arXiv: 1511.07376v1.
Apple Inc. OpenCL reference guide [EB/OL].(2017-08-17)[2017-08-22]. https: ∥www.khronos.org/files/opencl22-reference-guide.pdf.
SONG I, KIM H J, JEON P B. Deep learning for real-time robust facial expression recognition on a smartphone [C]∥Proceedings of the IEEE International Conference on Consumer Electronics. Piscataway, NJ, USA: IEEE, 2014: 564-567.
FARABET C, MARTINI B, AKSELROD P, et al. Hardware accelerated convolutional neural networks for synthetic vision systems [J]. IEEE International Symposium on Circuits and Systems, 2010, 54(3): 257-260.
CUI Jiyue, MEI Kuizhi, LIU Dongdong, et al. Construction of embedded Mali GPU simulator for OpenCL [J]. Journal of Xi'an Jiaotong University, 2015, 49(2): 20-24.