the demand for computing resources and memory bandwidth in deep-level and large-scale deep learning network models is increasing exponentially. Traditional industry solution CPU+GPU is not suitable to the prevalent scenarios of mobile embedded applications. To deal with this problem
we proposed a design of convolutional neural network co-processor based on FPGA programmable logic device. This solution focuses on high compatibility. It has programmability and is compatible with a variety of network models to achieve hardware acceleration. It also has scalability to allow multi-core expansion within the range of hardware resources to achieve double
performance. The design of convolutional operation module focuses on hardware parallelism and data reusability
which improves the utilization of hardware resources and computing efficiency. Rationally configured multi-level buffer structure reduces the co-processor's occupancy rate of external memory's read/write frequency and bandwidth
improves the internal communication efficiency of the module. The experimental results on the XILINX VC707 evaluation board show that the accuracy of the test set is 99%
the CIFAR-10 can achieve 80%
and the peak computing capability is 5.511×10
10
s
-1
the overall performance is approximately twice that of the general-purpose processor of Intel Xeno E5-2640 V4 server. Moreover
the processing performance of our design reaches the current mainstream level of FPGA solutions.
关键词
Keywords
references
RUMELHART D E, HINTON G E, WILLIAMS R J. Learning representations by back-propagating errors [J]. Nature, 1986, 323(6088): 533-536.
DALY D, FUJINO L. ISSCC 2017: intelligent chips for a smart world [J]. IEEE Solid-State Circuits Magazine, 2016, 8(4): 92-93.
FRIEDMAN D. Hardware approaches to machine learning and inference [C]∥2018 IEEE International Solid-State Circuits Conference. Piscataway, NJ, USA: IEEE, 2018: 2376-8606.
CHEN T, DU Z, SUN N, et al. DianNao: a small-footprint high-throughput accelerator for ubiquitous machine-learning [J]. ACM SIGPLAN Notices, 2014, 49(4): 269-284.
LU Qi. Baidu's brain is the core of Baidu's AI platform Smart Cloud has the opportunity to subvert the cloud market [J]. China Computer Communication, 2017(14): 1-2.
LU Hongtao, ZHANG Qinchuan. Applications of deep convolutional neural network in computer vision [J]. Journal of Data Acquisition and Processing, 2016, 31(1): 1-17.
ANTHIMOPOULOS M, CHRISTODOULIDIS S, EBNER L, et al. Lung pattern classification for interstitial lung diseases using a deep convolutional neural network [J]. IEEE Transactions on Medical Imaging, 2016, 35(5): 1207.
KRIZHEVSKY A, SUTSKEVER I, HINTON G E. ImageNet classification with deep convolutional neural networks [C]∥Advances in Neural Information Processing Systems. Piscataway, NJ, USA: IEEE, 2012: 1097-1105.
LONG J, SHELHAMER E, DARRELL T. Fully convolutional networks for semantic segmentation [C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2015: 3431-3440.
CHEN Jianying, YANG Xianze, ZHANG Nan. Research on multi-level structure of cache information in large-scale distributed system [J]. Journal of Southwest University for Nationalities, 2012, 38(3): 457-460.
ZHANG C, LI P, SUN G, et al. Optimizing FPGA-based accelerator design for deep convolutional neural networks [C]∥Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays. New York, USA: ACM, 2015: 161-170.
WANG D, AN J, XU K. PipeCNN: an OpenCL-based FPGA accelerator for large-scale convolution neuron networks [EB/OL]. [2018-03-16]. http:∥pdfs. semanticscholar.org/8d6d/df21989e9b5bd15e4bf f972e2370d5dc47d2.pdf.
QIU J, WANG J, YAO S, et al. Going deeper with embedded FPGA platform for convolutional neural network [C]∥Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays. New York, USA: ACM, 2016: 26-35.