西安建筑科技大学信息与控制工程学院,西安,710055
网络首发:2020-05-10,
纸质出版:2020
移动端阅览
孟月波, 纪拓, 刘光辉, 等. 编码-解码多尺度卷积神经网络人群计数方法[J]. 西安交通大学学报, 2020,54(5):149-157.
Encoding-Decoding Multi-Scale Convolutional Neural Network for Crowd Counting[J]. 2020, 54(5): 149-157.
孟月波, 纪拓, 刘光辉, 等. 编码-解码多尺度卷积神经网络人群计数方法[J]. 西安交通大学学报, 2020,54(5):149-157. DOI: 10.7652/xjtuxb202005020.
Encoding-Decoding Multi-Scale Convolutional Neural Network for Crowd Counting[J]. 2020, 54(5): 149-157. DOI: 10.7652/xjtuxb202005020.
针对基于多列卷积神经网络的人群计数方法存在的多尺度特征信息丢失、融合不佳以及密度图质量不高等问题
提出了一种编码-解码结构的多尺度卷积神经网络人群计数方法。编码器采用多列卷积捕获多尺度特征
通过空洞空间金字塔池化扩大感受野并减少参数量
保留尺度特征和图像的上下文信息; 解码器对编码器输出进行上采样
实现高层语义信息和编码器前端低层特征信息有效融合
从而提升了密度图的输出质量。为增强网络对计数的敏感性
在以往像素空间损失的基础上考虑了计数误差
提出了一种新型损失函数。采用Shanghai Tech、Mall以及自建数据集进行了对比实验
结果表明:与之前最优方法相比
所提方法在Shanghai Tech数据集Part_A部分的平均绝对误差和均方误差分别降低了8.3%和21.3%
Part_B部分分别降低了12.9%和12.0%
Mall数据集分别降低了15.1%和23.8%
自建数据集分别降低了13.5%和7.1%; 在不同人群场景下
所提方法的人群计数准确性和鲁棒性均优于其他对比方法的。
Aiming at the problems of multi-scale feature information loss
poor fusion and low quality of density map in the crowd counting method based on multi-column convolutional neural network
a new crowd counting method is proposed based on encoding-decoding multi-scale convolutional neural network. The encoder part adopts multi-column convolution to capture multi-scale features
expands the receptive field and reduces the amount of calculation via atrous space pyramid pooling
and retains the multi-scale feature and the context information of the image. The decoder part upsamples the encoder output to achieve effective fusion of the features with rich high-level semantic information and the features with rich low-level detail information to improve the output quality of the density map. To enhance the sensitivity of the network to counting
a new loss function is proposed by considering the previous pixel space loss and the counting error. Contrast experiments with previous methods on Sha
CHEN K, LOY C C, GONG S, et al. Feature mining for localised crowd counting [C]∥Proceedings of the 2012 British Machine Vision Conference(BMVC). Guildford, UK: BMVA Press, 2012: 3-13.
CAI Zebin, YU Zhu Liang, LIU Hao. Counting people in crowded scenes by video analyzing [C]∥Proceedings of the 2014 IEEE 9th Conference on Industrial Electronics and Applications(ICIEA). Piscataway, NJ, USA: IEEE, 2014: 1841-1845.
CHEN T Y, CHEN C H, WANG D J, et al. A people counting system based on face-detection [C]∥Proceedings of the 4th International Conference on Genetic and Evolutionary Computing(ICGEC). Piscataway, NJ, USA: IEEE, 2011: 699-702.
MARANA A N, VELASTIN S A, COSTA L F, et al. Estimation of crowd density using image processing [C]∥Proceedings of the 1997 IEE Colloquium on Image Processing for Security Applications. Stevenage, UK: IEE, 1997: 11-1-11-8.
KILAMB P, RIBNICK E, JOSHI A J, et al. Estimating pedestrian counts in groups [J]. Computer Vision and Image Understanding, 2008, 110(1): 43-59.
CHEN K, GONG S, XIANG T, et al. Cumulative attribute space for age and crowd density estimation [C]∥Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2013: 2467-2474.
CHANGE L C, GONG S, XIANG T. From semi-supervised to transfer counting of crowds [C]∥Proceedings of the 2013 IEEE International Conference on Computer Vision(CVPR). Piscataway, NJ, USA: IEEE, 2013: 2256-2263.
CHO S Y, CHOW T W S, LEUNG C T. A neural-based crowd estimation by hybrid global learning algorithm [J]. IEEE Transactions on Cybernetics, 1999, 29(4): 535-541.
REN S, HE K, GIRSHICK R, et al. Faster R-CNN: towards real-time object detection with region proposal networks [C]∥Proceedings of the 2015 Annual Conference on Neural Information Processing Systems(NIPS). Vancouver, Canada: NIPS, 2015: 91-99.
曹玉良, 明廷锋, 贺国, 等. 基于深度学习的离心泵空化状态识别 [J]. 西安交通大学学报, 2017, 51(11): 165-172.
CAO Yuliang, MING Yanfeng, HE Guo, et al. Artificial recognition of centrifugal pump cavitation status based on deep learning [J]. Journal of Xi’an Jiaotong University, 2017, 51(11): 165-172.
常亮, 邓小明, 周明全, 等. 图像理解中的卷积神经网络 [J]. 自动化学报, 2016, 42(9): 1300-1312.
CHANG Liang, DENG Xiaoming, ZHOU Mingquan, et al. Convolutional neural networks in image understanding [J]. Journal of Automation, 2016, 42(9): 1300-1312.
KRIZHEVSKY A, SUTSKEVER I, HINTON G E. ImageNet classification with deep convolutional neural networks [J]. Communications of the ACM, 2017, 60(6): 84-90.
WANG C, ZHANG H, YANG L, et al. Deep people counting in extremely dense crowds [C]∥Proceedings of the 23rd ACM International Conference on Multimedia. New York, USA: ACM, 2015: 1299-1302.
ZHANG C, LI H, WANG X, et al. Cross-scene crowd counting via deep convolutional neural networks [C]∥Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2015: 833-841.
ZHANG Y, ZHOU D, CHEN S, et al. Single-image crowd counting via multi-column convolutional neural network [C]∥Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2016: 589-597.
SIMONYAN K, ZISSERMAN A. Very deep convolutional networks for large-scale image recognition [C/OL]∥Proceedings of the 3rd International Conference on Learning Representations(ICLR). London, UK: ICLR, 2015. [2019-07-01]. https:∥arxiv.org/pdf/1409.1556.pdf.
SAM D B, SURYA S, BABU R V. Switching convolutional neural network for crowd counting [C]∥Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2017: 4031-4039.
ZENG L, XU X, CAI B, et al. Multi-scale convolutional neural networks for crowd counting [C]∥Proceedings of the 2017 IEEE International Conference on Image Processing(ICIP). Piscataway, NJ, USA: IEEE, 2017: 465-469.
LEMPITSKY V, ZISSERMAN A. Learning to count objects in images [C]∥Proceedings of the Advances in Neural Information Processing Systems. Vancouver, Canada: NIPS, 2010: 1324-1332.
JOSEPH R, SANTOSH D, ROSS G. You only look once: unified, real-time object detection [C]∥Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2016: 779-788.
REDMON J, FARHADI A. YOLO9000: better, faster, stronger [C]∥Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2017: 7263-7271.
CHEN L C, ZHU Y, PAPANDREOU G, et al. Encoder-decoder with atrous separable convolution for semantic image segmentation [C]∥Proceedings of the European Conference on Computer Vision(ECCV). Berlin, Germany: Springer, 2018: 801-818.
张焯林, 赵建伟, 曹飞龙. 构建带空洞卷积的深度神经网络重建高分辨率图像 [J]. 模式识别与人工智能, 2019, 32(3): 259-267.
ZHANG Zhuolin, ZHAO Jianwei, CAO Feilong. Building deep neural networks with dilated convolutions to reconstruct high-resolution image [J]. Pattern Recognition and Artificial Intelligence, 2019, 32(3): 259-267.
HE K, ZHANG X, REN S, et al. Spatial pyramid pooling in deep convolutional networks for visual recognition [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2015, 37(9): 1904-1916.
时增林, 叶阳东, 吴云鹏, 等. 基于序的空间金字塔池化网络的人群计数方法 [J]. 自动化学报, 2016, 42(6): 866-874.
SHI Zenglin, YE Yangdong, WU Yunpeng, et al. Crowd counting using rank-based spatial pyramid pooling network [J]. Journal of Automation, 2016, 42(6): 866-874.
SINDAGI V A, PATEL V M. A survey of recent advances in CNN-based single image crowd counting and density estimation [J]. Pattern Recognition Letters, 2018, 107: 3-16.
0
浏览量
4
下载量
7
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621