1. 西安交通大学电子与信息学部,西安,710049
2. 上海机电工程研究所,上海,201109
: 2022-01-18。作者简介: 李悄(1996—),男,硕士生
李垚辰(通信作者),男,副教授,博士生导师。基金项目: 国家自然科学基金资助项目(61803298)
网络首发:2022-09-10,
纸质出版:2022
移动端阅览
李悄, 李垚辰, 张玉龙, 等. 采用稀疏3D卷积的单阶段点云三维目标检测方法[J]. 西安交通大学学报, 2022,56(9):112-122.
LI Qiao, LI Yaochen, ZHANG Yulong, et al. A Single-Stage Point Cloud 3D Object Detection Method Using Sparse 3D Convolution[J]. 2022, 56(9): 112-122.
李悄, 李垚辰, 张玉龙, 等. 采用稀疏3D卷积的单阶段点云三维目标检测方法[J]. 西安交通大学学报, 2022,56(9):112-122. DOI: 10.7652/xjtuxb202209012.
LI Qiao, LI Yaochen, ZHANG Yulong, et al. A Single-Stage Point Cloud 3D Object Detection Method Using Sparse 3D Convolution[J]. 2022, 56(9): 112-122. DOI: 10.7652/xjtuxb202209012.
针对点云体素化的三维目标检测方法中点云的特征提取能力不足的问题
将三维目标检测方法SECOND作为基准网络
提出一种采用稀疏3D卷积的单阶段点云三维目标检测(Reinforced SECOND)方法。首先
改进点云分组方式形成鲁棒的抽离体素特征的体素特征编码网络; 其次
为增强体素中对检测任务有显著贡献的关键特征
同时抑制不相关噪声特征
把堆叠三重注意力机制引入体素特征编码网络; 然后
提出残差稀疏卷积单元
设计了残差稀疏卷积中间网络
提高了该网络层的特征提取能力
保留了更多的原始特征信息; 最后
把提出的空间语义特征融合(SSFF)模块引入区域建议网络
自适应地融合低级空间特征和高级抽象语义特征
进一步提高了模型特征的表达能力。在KITTI开源数据集上的实验结果表明:与之前许多基于网格及基于点的方法相比
所提方法显著提高了三维目标检测性能; 与基准网络相比
采用所提方法对KITTI测试集car类和cyclist类进行检测
中等难度级别下的3D检测精度分别提高了5.85%和8.9%
困难难度级别下的3D检测精度分别提高了8.54%和8.53%。
In view of the insufficient ability to extract features of point clouds for 3D object detection methods based on point cloud voxelization
this paper proposes a single-stage point cloud 3D object detection method(Reinforced SECOND)based on sparse 3D convolution using SECOND as the baseline. Firstly
the point cloud grouping method is improved herein to form a robust voxel feature encoding network that extracts voxel features. Secondly
to further enhance the key features in voxels that significantly contribute to the detection task while suppressing irrelevant and noisy features
this paper introduces the stacked triple attention mechanism into the voxel feature encoding network. Then
this paper proposes the residual sparse 3D convolution unit and designs the residual sparse convolution intermediate network
which improves the feature extraction ability of this network layer and retains more original feature information. Finally
this paper introduces the proposed spatial-semantic feature fusion(SSFF)module into the region proposal network
which adaptively fuses low-level spatial features and high-level abstract semantic features
further improving the expressiveness of model features. Experimental results on the KITTI open-source dataset demonstrate the proposed method significantly improves 3D object detection performance compared to many previous grid-based and point-based methods. Compared with the benchmark network of this method
3D car and cyclist detection in the KITTI test set are improved by 5.85% and 8.9% on the moderate mAP and by 8.54% and 8.53% on the hard mAP.
陈科圻, 朱志亮, 邓小明, 等. 多尺度目标检测的深度学习研究综述 [J]. 软件学报, 2021, 32(4): 1201-1227.
CHEN Keqi, ZHU Zhiliang, DENG Xiaoming, et al. Deep learning for multi-scale object detection: a survey [J]. Journal of Software, 2021, 32(4): 1201-1227.
张帆, 赵世坤, 袁操, 等. 人脸识别反欺诈研究进展 [J]. 软件学报, 2022, 33(7): 2204-2240.
ZHANG Fan, ZHAO Shikun, YUAN Cao, et al. Recent progress of face anti-spoofing [J]. Journal of Software, 2022, 33(7): 2204-2240.
陈晋音, 沈诗婧, 苏蒙蒙, 等. 车牌识别系统的黑盒对抗攻击 [J]. 自动化学报, 2021, 47(1): 121-135.
CHEN Jinyin, SHEN Shijing, SU Mengmeng, et al. Black-box adversarial attack on license plate recognition system [J]. Acta Automatica Sinica, 2021, 47(1): 121-135.
孟琭, 杨旭. 目标跟踪算法综述 [J]. 自动化学报, 2019, 45(7): 1244-1260.
MENG Lu, YANG Xu. A survey of object tracking algorithms [J]. Acta Automatica Sinica, 2019, 45(7): 1244-1260.
田永林, 沈宇, 李强, 等. 平行点云: 虚实互动的点云生成与三维模型进化方法 [J]. 自动化学报, 2020, 46(12): 2572-2582.
TIAN Yonglin, SHEN Yu, LI Qiang, et al. Parallel point clouds: point clouds generation and 3D model evolution via virtual-real interaction [J]. Acta Automatica Sinica, 2020, 46(12): 2572-2582.
QI C R, LIU Wei, WU Chenxia, et al. Frustum PointNets for 3D object detection from RGB-D data [C]∥2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 918-927.
SHI Shaoshuai, WANG Xiaogang, LI Hongsheng. PointRCNN: 3D object proposal generation and detection from point cloud [C]∥2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2019: 770-779.
YANG Zetong, SUN Yanan, LIU Shu, et al. 3DSSD: point-based 3D single stage object detector [C]∥2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2020: 11037-11045.
YANG Zetong, SUN Yanan, LIU Shu, et al. STD: sparse-to-dense 3D object detector for point cloud [C]∥2019 IEEE/CVF International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2019: 1951-1960.
QI C R, LITANY O, HE Kaiming, et al. Deep hough voting for 3D object detection in point clouds [C]∥2019 IEEE/CVF International Conference on Computer Vision(ICCV). Piscataway, NJ, USA: IEEE, 2019: 9276-9285.
CHARLES R Q, SU Hao, KAICHUN Mo, et al. PointNet: deep learning on point sets for 3D classification and segmentation [C]∥2017 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2017: 77-85.
QI C R, YI Li, SU Hao, et al. PointNet++: deep hierarchical feature learning on point sets in a metric space [C]∥Proceedings of the 31st International Conference on Neural Information Processing Systems. Red Hook, NY, USA: Curran Associates Inc., 2017: 5105-5114.
SHI Shaoshuai, WANG Zhe, WANG Xiaogang, et al. Part-A2 net: 3D part-aware and aggregation neural network for object detection from point cloud [EB/OL]. [2021-12-09]. https:∥doi.org/10.48550/arXiv. 1907.03670.
SINDAGI V A, ZHOU Yin, TUZEL O. MVX-net: multimodal VoxelNet for 3D object detection [C]∥2019 International Conference on Robotics and Automation(ICRA). Piscataway, NJ, USA: IEEE, 2019: 7276-7282.
YAN Yan, MAO Yuxing, LI Bo. SECOND: sparsely embedded convolutional detection [J]. Sensors, 2018, 18(10): 3337.
ZHOU Yin, TUZEL O. VoxelNet: end-to-end learning for point cloud based 3D object detection [C]∥2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 4490-4499.
LANG A H, VORA S, CAESAR H, et al. PointPillars: fast encoders for object detection from point clouds [C]∥2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2019: 12689-12697.
SIMON M, MILZ S, AMENDE K, et al. Complex-YOLO: an Euler-region-proposal for real-time 3D object detection on point clouds [C]∥Computer Vision: ECCV 2018 Workshops. Cham, Switzerland: Springer International Publishing, 2019: 197-209.
YANG Bin, LUO Wenjie, URTASUN R. PIXOR: real-time 3D object detection from point clouds [C]∥2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 7652-7660.
REN Shaoqing, HE Kaiming, GIRSHICK R, et al. Faster R-CNN: towards real-time object detection with region proposal networks [C]∥Proceedings of the 28th International Conference on Neural Information Processing Systems: Volume 1. Cambridge, MA, USA: MIT Press, 2015: 91-99.
GRAHAM B. Sparse 3D convolutional neural networks [EB/OL]. [2021-12-09]. https:∥doi.org/10.48550/arXiv.1505.02890.
GRAHAM B. VAN DER MAATEN L. Submanifold sparse convolutional networks [EB/OL]. [2021-12-09]. https:∥doi.org/10.48550/arXiv.1706.01307.
LIU Zhe, ZHAO Xin, HUANG Tengteng, et al. TANet: robust 3D object detection from point clouds with triple attention [C]∥Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto, CA, USA: AAAI Press, 2020: 11677-11684.
HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition [C]∥2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway, NJ, USA: IEEE, 2016: 770-778.
KU J, MOZIFIAN M, LEE J, et al. Joint 3D proposal generation and object detection from view aggregation [C]∥2018 IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS). Piscataway, NJ, USA: IEEE, 2018: 1-8.
0
浏览量
13
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621