1. 西安工业大学电子信息工程学院,西安,710021
2. 西安翔迅科技有限责任公司,西安,710068
3. 西安工业大学发展规划处,西安,710021
4. 陕西航天技术应用研究院有限公司,西安,710100
: 2022-05-20。作者简介: 李晓艳(1982—),女,副教授
王鹏(通信作者),男,教授。基金项目: 国家自然科学基金资助项目(62171360)
网络首发:2022-10-10,
纸质出版:2022
移动端阅览
李晓艳, 符惠桐, 牛文涛, 等. 基于深度学习的多模态行人检测算法[J]. 西安交通大学学报, 2022,56(10):61-70.
Multi-Modal Pedestrian Detection Algorithm Based on Deep Learning[J]. 2022, 56(10): 61-70.
李晓艳, 符惠桐, 牛文涛, 等. 基于深度学习的多模态行人检测算法[J]. 西安交通大学学报, 2022,56(10):61-70. DOI: 10.7652/xjtuxb202210006.
Multi-Modal Pedestrian Detection Algorithm Based on Deep Learning[J]. 2022, 56(10): 61-70. DOI: 10.7652/xjtuxb202210006.
针对全天候工作的多模态行人检测算法体积大、运算量高、效率不足的问题
提出一种基于深度学习MBNet算法搭建的轻量级多模态行人检测算法(G-MBNet)。采用ResNet18算法并结合跨阶段链接的思想搭建CSP-ResNet18轻量级特征提取网络
以保证检测算法精度; 引入轻量级高效通道注意力(ECA)模块来提升特征提取网络对重要特征的关注能力
在引入极少参数的情况下提升算法的检测精度; 通过引入轻量级Ghost卷积模块来重构MBNet算法的特征提取网络
在保证特征提取性能的情况下进一步降低算法的参数与体积
提升算法的检测速度。采用所提的G-MBNet算法在KAIST行人数据集进行测试
实验结果表明:G-MBNet算法大小是原始算法的32.33%
参数量是原始算法的37.81%
检测速度是原始算法的1.53倍; G-MBNet算法可在保证行人识别精度的情况下有效提升检测速度。
This paper proposes a lightweight multi-modal pedestrian detection algorithm based on deep learning MBNet algorithm to address the problem of 24/7 multi-modal pedestrian detection algorithm having a large size
high computation load and insufficient efficiency. First
the CSP-ResNet18 lightweight feature extraction network is built using the ResNet18 algorithm combined with the idea of cross-stage linking to ensure accuracy of the detection algorithm; then
the lightweight efficient channel attention module is introduced to enhance the feature extraction network's ability to focus on important features
which can improve the detection accuracy of the algorithm by introducing very few parameters; finally
the feature extraction network of MBNet algorithm is reconstructed by introducing a lightweight Ghost convolution module
which can further reduce the parameters required and the size of the algorithm and improve its detection efficiency while ensuring the feature extraction performance. The proposed G-MBNet algorithm is tested on the KAIST pedestrian dataset
and the experimental results show that the size of G-MBNet algorithm is 32.33% that of the original algorithm; the number of parameters is 37.81% that of the original algorithm; the detection speed is 1.53 times that of the original algorithm. The experiments verify that G-MBNet algorithm can effectively improve the detection speed while ensuring the pedestrian recognition accuracy.
童靖然, 毛力, 孙俊. 特征金字塔融合的多模态行人检测算法 [J]. 计算机工程与应用, 2019, 55(19): 214-222.
TONG Jingran, MAO Li, SUN Jun. Multimodal pedestrian detection algorithm based on fusion feature pyramids [J]. Computer Engineering and Applications, 2019, 55(19): 214-222.
符惠桐, 王鹏, 李晓艳, 等. 面向移动目标识别的轻量化网络模型 [J]. 西安交通大学学报, 2021, 55(7): 124-131.
FU Huitong, WANG Peng, LI Xiaoyan, et al. Lightweight network model for moving object recognition [J]. Journal of Xi'an Jiaotong University, 2021, 55(7): 124-131.
REN Shaoqing, HE Kaiming, GIRSHICK R, et al. Faster R-CNN: towards real-time object detection with region proposal networks [C]//Proceedings of the 2015 Conference on Advances in Neural Information Processing Systems. Vancouver, Canada: NIPS, 2015: 91-99.
LIU Wei, ANGUELOV D, ERHAN D, et al. SSD: single shot multiBox detector [C]//Computer Vision-ECCV 2016. Cham, Germany: Springer International Publishing, 2016: 21-37.
REDMON J, DIVVALA S, GIRSHICK R, et al. You only look once: unified, real-time object detection [C]//2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2016: 779-788.
乔梦雨, 王鹏, 吴娇, 等. 面向陆战场目标识别的轻量级卷积神经网络 [J]. 计算机科学, 2020, 47(5): 161-165.
QIAO Mengyu, WANG Peng, WU Jiao, et al. Lightweight convolutional neural networks for land battle target recognition [J]. Computer Science, 2020, 47(5): 161-165.
程腾, 孙磊, 侯登超, 等. 基于特征融合的多层次多模态目标检测 [J]. 汽车工程, 2021, 43(11): 1602-1610.
CHENG Teng, SUN Lei, HOU Dengchao, et al. Multi-level and multi-modal target detection based on feature fusion [J]. Automotive Engineering, 2021, 43(11): 1602-1610.
HWANG S, PARK J, KIM N, et al. Multispectral pedestrian detection: benchmark dataset and baseline [C]//2015 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2015: 1037-1045.
CHOI Y, KIM N, HWANG S, et al. KAIST multi-spectral day/night data set for autonomous and assisted driving [J]. IEEE Transactions on Intelligent Transportation Systems, 2018, 19(3): 934-948.
LIU Jingjing, ZHANG Shaoting, WANG Shu, et al. Multispectral deep neural networks for pedestrian detection [EB/OL].[2022-03-01]. https://arxiv.org/abs/1611.02644.
KÖNIG D, ADAM M, JARVERS C, et al. Fully convolutional region proposal networks for multispectral person detection [C]//2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops(CVPRW). Piscataway, NJ, USA: IEEE, 2017: 243-250.
LI Chengyang, SONG Dan, TONG Ruofeng, et al. Illumination-aware faster R-CNN for robust multispectral pedestrian detection [J]. Pattern Recognition, 2019, 85: 161-171.
HE Kaiming, GKIOXARI G, DOLLÁR P, et al. Mask R-CNN [C]//2017 IEEE International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2017: 2980-2988.
GIRSHICK R. Fast R-CNN [C]//2015 IEEE International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2015: 1440-1448.
ZHANG Lu, LIU Zhiyong, ZHANG Shifeng, et al. Cross-modality interactive attention network for multispectral pedestrian detection [J]. Information Fusion, 2019, 50: 20-29.
ZHOU Kailai, CHEN Linsen, CAO Xun. Improving multispectral pedestrian detection by addressing modality imbalance problems [C]//Computer Vision-ECCV 2020. Cham, Germany: Springer International Publishing, 2020: 787-803.
HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition [C]//2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2016: 770-778.
WANG C Y, MARK LIAO H Y, WU Y H, et al. CSPNet: a new backbone that can enhance learning capability of CNN [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops(CVPRW). Piscataway, NJ, USA: IEEE, 2020: 1571-1580.
WANG Qilong, WU Banggu, ZHU Pengfei, et al. ECA-Net: efficient channel attention for deep convolutional neural networks [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2020: 11531-11539.
HAN Kai, WANG Yunhe, TIAN Qi, et al. GhostNet: more features from cheap operations [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2020: 1577-1586.
HU Jie, SHEN Li, SUN Gang. Squeeze-and-excitation networks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 7132-7141.
LIN T Y, MAIRE M, BELONGIE S, et al. Microsoft COCO: common objects in context [C]//Computer Vision-ECCV 2014. Cham, Germany: Springer International Publishing, 2014: 740-755.
WOO S, PARK J, LEE J Y, et al. CBAM: convolutional block attention module [C]//Computer Vision-ECCV 2018. Cham, Germany: Springer International Publishing, 2018: 3-19.
HOWARD A G, ZHU Menglong, CHEN Bo, et al. MobileNets: efficient convolutional neural networks for mobile vision applications [EB/OL].[2022-03-01]. http://arxiv.org/abs/1704.04861.
SANDLER M, HOWARD A, ZHU Menglong, et al. MobileNetV2: inverted residuals and linear bottlenecks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 4510-4520.
0
浏览量
5
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621