This paper proposes a lightweight multi-modal pedestrian detection algorithm based on deep learning MBNet algorithm to address the problem of 24/7 multi-modal pedestrian detection algorithm having a large size
high computation load and insufficient efficiency. First
the CSP-ResNet18 lightweight feature extraction network is built using the ResNet18 algorithm combined with the idea of cross-stage linking to ensure accuracy of the detection algorithm; then
the lightweight efficient channel attention module is introduced to enhance the feature extraction network's ability to focus on important features
which can improve the detection accuracy of the algorithm by introducing very few parameters; finally
the feature extraction network of MBNet algorithm is reconstructed by introducing a lightweight Ghost convolution module
which can further reduce the parameters required and the size of the algorithm and improve its detection efficiency while ensuring the feature extraction performance. The proposed G-MBNet algorithm is tested on the KAIST pedestrian dataset
and the experimental results show that the size of G-MBNet algorithm is 32.33% that of the original algorithm; the number of parameters is 37.81% that of the original algorithm; the detection speed is 1.53 times that of the original algorithm. The experiments verify that G-MBNet algorithm can effectively improve the detection speed while ensuring the pedestrian recognition accuracy.
TONG Jingran, MAO Li, SUN Jun. Multimodal pedestrian detection algorithm based on fusion feature pyramids [J]. Computer Engineering and Applications, 2019, 55(19): 214-222.
FU Huitong, WANG Peng, LI Xiaoyan, et al. Lightweight network model for moving object recognition [J]. Journal of Xi'an Jiaotong University, 2021, 55(7): 124-131.
REN Shaoqing, HE Kaiming, GIRSHICK R, et al. Faster R-CNN: towards real-time object detection with region proposal networks [C]//Proceedings of the 2015 Conference on Advances in Neural Information Processing Systems. Vancouver, Canada: NIPS, 2015: 91-99.
LIU Wei, ANGUELOV D, ERHAN D, et al. SSD: single shot multiBox detector [C]//Computer Vision-ECCV 2016. Cham, Germany: Springer International Publishing, 2016: 21-37.
REDMON J, DIVVALA S, GIRSHICK R, et al. You only look once: unified, real-time object detection [C]//2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2016: 779-788.
CHENG Teng, SUN Lei, HOU Dengchao, et al. Multi-level and multi-modal target detection based on feature fusion [J]. Automotive Engineering, 2021, 43(11): 1602-1610.
HWANG S, PARK J, KIM N, et al. Multispectral pedestrian detection: benchmark dataset and baseline [C]//2015 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2015: 1037-1045.
CHOI Y, KIM N, HWANG S, et al. KAIST multi-spectral day/night data set for autonomous and assisted driving [J]. IEEE Transactions on Intelligent Transportation Systems, 2018, 19(3): 934-948.
LIU Jingjing, ZHANG Shaoting, WANG Shu, et al. Multispectral deep neural networks for pedestrian detection [EB/OL].[2022-03-01]. https://arxiv.org/abs/1611.02644.
KÖNIG D, ADAM M, JARVERS C, et al. Fully convolutional region proposal networks for multispectral person detection [C]//2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops(CVPRW). Piscataway, NJ, USA: IEEE, 2017: 243-250.
LI Chengyang, SONG Dan, TONG Ruofeng, et al. Illumination-aware faster R-CNN for robust multispectral pedestrian detection [J]. Pattern Recognition, 2019, 85: 161-171.
HE Kaiming, GKIOXARI G, DOLLÁR P, et al. Mask R-CNN [C]//2017 IEEE International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2017: 2980-2988.
GIRSHICK R. Fast R-CNN [C]//2015 IEEE International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2015: 1440-1448.
ZHANG Lu, LIU Zhiyong, ZHANG Shifeng, et al. Cross-modality interactive attention network for multispectral pedestrian detection [J]. Information Fusion, 2019, 50: 20-29.
ZHOU Kailai, CHEN Linsen, CAO Xun. Improving multispectral pedestrian detection by addressing modality imbalance problems [C]//Computer Vision-ECCV 2020. Cham, Germany: Springer International Publishing, 2020: 787-803.
HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition [C]//2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2016: 770-778.
WANG C Y, MARK LIAO H Y, WU Y H, et al. CSPNet: a new backbone that can enhance learning capability of CNN [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops(CVPRW). Piscataway, NJ, USA: IEEE, 2020: 1571-1580.
WANG Qilong, WU Banggu, ZHU Pengfei, et al. ECA-Net: efficient channel attention for deep convolutional neural networks [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2020: 11531-11539.
HAN Kai, WANG Yunhe, TIAN Qi, et al. GhostNet: more features from cheap operations [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2020: 1577-1586.
HU Jie, SHEN Li, SUN Gang. Squeeze-and-excitation networks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 7132-7141.
LIN T Y, MAIRE M, BELONGIE S, et al. Microsoft COCO: common objects in context [C]//Computer Vision-ECCV 2014. Cham, Germany: Springer International Publishing, 2014: 740-755.
WOO S, PARK J, LEE J Y, et al. CBAM: convolutional block attention module [C]//Computer Vision-ECCV 2018. Cham, Germany: Springer International Publishing, 2018: 3-19.
HOWARD A G, ZHU Menglong, CHEN Bo, et al. MobileNets: efficient convolutional neural networks for mobile vision applications [EB/OL].[2022-03-01]. http://arxiv.org/abs/1704.04861.
SANDLER M, HOWARD A, ZHU Menglong, et al. MobileNetV2: inverted residuals and linear bottlenecks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 4510-4520.