LI Qiao, LI Yaochen, ZHANG Yulong, et al. A Single-Stage Point Cloud 3D Object Detection Method Using Sparse 3D Convolution[J]. 2022, 56(9): 112-122.
DOI:
LI Qiao, LI Yaochen, ZHANG Yulong, et al. A Single-Stage Point Cloud 3D Object Detection Method Using Sparse 3D Convolution[J]. 2022, 56(9): 112-122.DOI: 10.7652/xjtuxb202209012.
A Single-Stage Point Cloud 3D Object Detection Method Using Sparse 3D Convolution
In view of the insufficient ability to extract features of point clouds for 3D object detection methods based on point cloud voxelization
this paper proposes a single-stage point cloud 3D object detection method(Reinforced SECOND)based on sparse 3D convolution using SECOND as the baseline. Firstly
the point cloud grouping method is improved herein to form a robust voxel feature encoding network that extracts voxel features. Secondly
to further enhance the key features in voxels that significantly contribute to the detection task while suppressing irrelevant and noisy features
this paper introduces the stacked triple attention mechanism into the voxel feature encoding network. Then
this paper proposes the residual sparse 3D convolution unit and designs the residual sparse convolution intermediate network
which improves the feature extraction ability of this network layer and retains more original feature information. Finally
this paper introduces the proposed spatial-semantic feature fusion(SSFF)module into the region proposal network
which adaptively fuses low-level spatial features and high-level abstract semantic features
further improving the expressiveness of model features. Experimental results on the KITTI open-source dataset demonstrate the proposed method significantly improves 3D object detection performance compared to many previous grid-based and point-based methods. Compared with the benchmark network of this method
3D car and cyclist detection in the KITTI test set are improved by 5.85% and 8.9% on the moderate mAP and by 8.54% and 8.53% on the hard mAP.
CHEN Keqi, ZHU Zhiliang, DENG Xiaoming, et al. Deep learning for multi-scale object detection: a survey [J]. Journal of Software, 2021, 32(4): 1201-1227.
TIAN Yonglin, SHEN Yu, LI Qiang, et al. Parallel point clouds: point clouds generation and 3D model evolution via virtual-real interaction [J]. Acta Automatica Sinica, 2020, 46(12): 2572-2582.
QI C R, LIU Wei, WU Chenxia, et al. Frustum PointNets for 3D object detection from RGB-D data [C]∥2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 918-927.
SHI Shaoshuai, WANG Xiaogang, LI Hongsheng. PointRCNN: 3D object proposal generation and detection from point cloud [C]∥2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2019: 770-779.
YANG Zetong, SUN Yanan, LIU Shu, et al. 3DSSD: point-based 3D single stage object detector [C]∥2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2020: 11037-11045.
YANG Zetong, SUN Yanan, LIU Shu, et al. STD: sparse-to-dense 3D object detector for point cloud [C]∥2019 IEEE/CVF International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2019: 1951-1960.
QI C R, LITANY O, HE Kaiming, et al. Deep hough voting for 3D object detection in point clouds [C]∥2019 IEEE/CVF International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2019: 9276-9285.
CHARLES R Q, SU Hao, KAICHUN Mo, et al. PointNet: deep learning on point sets for 3D classification and segmentation [C]∥2017 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2017: 77-85.
QI C R, YI Li, SU Hao, et al. PointNet++: deep hierarchical feature learning on point sets in a metric space [C]∥Proceedings of the 31st International Conference on Neural Information Processing Systems. Red Hook, NY, USA: Curran Associates Inc., 2017: 5105-5114.
SHI Shaoshuai, WANG Zhe, WANG Xiaogang, et al. Part-A2 net: 3D part-aware and aggregation neural network for object detection from point cloud [EB/OL]. [2021-12-09]. https:∥doi.org/10.48550/arXiv. 1907.03670.
SINDAGI V A, ZHOU Yin, TUZEL O. MVX-net: multimodal VoxelNet for 3D object detection [C]∥2019 International Conference on Robotics and Automation(ICRA). Piscataway, NJ, USA: IEEE, 2019: 7276-7282.
YAN Yan, MAO Yuxing, LI Bo. SECOND: sparsely embedded convolutional detection [J]. Sensors, 2018, 18(10): 3337.
ZHOU Yin, TUZEL O. VoxelNet: end-to-end learning for point cloud based 3D object detection [C]∥2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 4490-4499.
LANG A H, VORA S, CAESAR H, et al. PointPillars: fast encoders for object detection from point clouds [C]∥2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2019: 12689-12697.
SIMON M, MILZ S, AMENDE K, et al. Complex-YOLO: an Euler-region-proposal for real-time 3D object detection on point clouds [C]∥Computer Vision: ECCV 2018 Workshops. Cham, Switzerland: Springer International Publishing, 2019: 197-209.
YANG Bin, LUO Wenjie, URTASUN R. PIXOR: real-time 3D object detection from point clouds [C]∥2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 7652-7660.
REN Shaoqing, HE Kaiming, GIRSHICK R, et al. Faster R-CNN: towards real-time object detection with region proposal networks [C]∥Proceedings of the 28th International Conference on Neural Information Processing Systems: Volume 1. Cambridge, MA, USA: MIT Press, 2015: 91-99.
GRAHAM B. Sparse 3D convolutional neural networks [EB/OL]. [2021-12-09]. https:∥doi.org/10.48550/arXiv.1505.02890.
GRAHAM B. VAN DER MAATEN L. Submanifold sparse convolutional networks [EB/OL]. [2021-12-09]. https:∥doi.org/10.48550/arXiv.1706.01307.
LIU Zhe, ZHAO Xin, HUANG Tengteng, et al. TANet: robust 3D object detection from point clouds with triple attention [C]∥Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto, CA, USA: AAAI Press, 2020: 11677-11684.
HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition [C]∥2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2016: 770-778.
KU J, MOZIFIAN M, LEE J, et al. Joint 3D proposal generation and object detection from view aggregation [C]∥2018 IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS). Piscataway, NJ, USA: IEEE, 2018: 1-8.