1.兰州理工大学自动化与电气工程学院,730050,兰州
2.兰州理工大学微电子现代产业学院,730050,兰州
3.人机混合增强智能全国重点实验室,710049,西安
4.西安交通大学人工智能与机器人研究所,710049,西安
5.西安交通大学第二附属医院超声医学科,710299,西安
收稿:2026-04-22,
修回:2026-06-24,
录用:2026-07-27,
移动端阅览
王宗顺, 李策, 王鹏程, 等. 面向自动驾驶的三维自适应池化点云分割网络[J]. 西安交通大学学报,2026.
WANG Zongshun, LI Ce, WANG Pengcheng, et al. 3D Adaptive Pooling Point Cloud Segmentation Network for Autonomous Driving[J]. JOURNAL OF XI’AN JIAOTONG UNIVERSITY,2026.
针对自动驾驶场景中车载激光雷达点云规模大、分布稀疏不均,现有点云Transformer网络存在邻域搜索开销大、体素注意力计算量高及关键目标局部几何细节易丢失的问题,提出了一种三维自适应池化点云分割网络(3DAPT-Net)。首先,在非重叠体素窗口内利用三维自适应池化压缩键值序列,聚合多尺度上下文,以降低大规模点云的注意力计算开销;其次,采用点注意力分支聚合点级特征,补充道路边界、行人轮廓及远距离稀疏目标的细粒度几何信息;最后,融合点级特征与体素特征,形成兼顾精度与效率的点云分割网络。该网络在SemanticKITTI和ShapeNet数据集上的平均交并比分别为67.5%和87.4%;与Point Cloud Transformer网络相比,推理延迟降低54.9%。结果表明,3DAPT-Net网络在分割精度与计算效率之间取得了较好的平衡。
In autonomous driving scenarios
vehicle-mounted LiDAR point clouds are characterized by large scale and sparse
uneven distribution. To address the problems of high neighborhood search overhead
heavy voxel attention computation load
and easy loss of local geometric details of key objects in existing point cloud Transformer networks
a 3D adaptive pooling point cloud segmentation network (3DAPT-Net) is proposed. First
3D adaptive pooling is utilized within non-overlapping voxel windows to compress key-value sequences and aggregate multiscale context
thereby reducing the attention computation overhead for large-scale point clouds. Second
a point attention branch is adopted to aggregate point-level features
supplementing the fine-grained geometric information of road boundaries
pedestrian contours
and distant sparse objects. Finally
point-level features and voxel features are fused to construct a point cloud segmentation network that balances both accuracy and efficiency. The proposed network achieves mean intersection over union (mIoU) scores of 67.5% and 87.4% on the SemanticKITTI and ShapeNet datasets
respectively; compared with the Point Cloud Transformer network
it has the inference latency reduced by 54.9%. The results indicate that 3DAPT-Net achieves a favorable balance between segmentation accuracy and computational efficiency.
Betsas T , Georgopoulos A , Doulamis A , et al . Deep learning on 3D semantic segmentation: a detailed review [J ] . Remote Sensing , 2025 , 17 ( 2 ): 298 .
Leng Wenfeng , Hu Chuan , Huang Yonghui , et al . KD-DiffSeg: knowledge distillation guided LiDAR-camera diffusion framework for 3D semantic segmentation [J ] . Expert Systems with Applications , 2026 , 322 : 132272 .
蔡子悦 , 袁振岳 , 庞明勇 . 深度学习的点云语义分割方法综述 [J ] . 计算机工程与应用 , 2025 , 61 ( 11 ): 22 - 30 .
Cai Ziyue , Yuan Zhenyue , Pang Mingyong . Survey on deep-learning-based point cloud semantic segmentation [J ] . Computer Engineering and Applications , 2025 , 61 ( 11 ): 22 - 30 .
Wang Zongshun , Li Ce , Ma Jialin , et al . Explicit geometric relationships under limited spatial reference points guide 3D visual grounding [J ] . Information Processing & Management , 2026 , 63 ( 6 ): 104720 .
Vaswani A , Shazeer N , Parmar N , et al . Attention is all you need [C ] // Proceedings of the 31st International Conference on Neural Information Processing Systems . Red Hook, NY, USA : Curran Associates Inc. , 2017 : 6000 - 6010 .
Charles R Q , Su Hao , Kaichun Mo , et al . PointNet: deep learning on point Sets for 3D classification and segmentation [C ] // 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2017 : 77 - 85 .
Li Wangkai , Li Zhaoyang , Pan Yuwen , et al . Adaptive augmentation-aware latent learning for robust LiDAR semantic segmentation [C ] // ICLR 2026 . New York, USA : ICLR , 2026 : 1 - 36 .
Eldar Y , Lindenbaum M , Porat M , et al . The farthest point strategy for progressive image sampling [J ] . IEEE Transactions on Image Processing , 1997 , 6 ( 9 ): 1305 - 1315 .
Cover Thomas M , Hart Peter E . Nearest neighbor pattern classification [J ] . IEEE Transactions on Information Theory , 1967 , 13 ( 1 ): 21 - 27 .
Behley J , Garbade M , Milioto A , et al . SemanticKITTI: a dataset for semantic scene understanding of LiDAR sequences [C ] // 2019 IEEE/CVF International Conference on Computer Vision (ICCV) . Piscataway, NJ, USA : IEEE , 2019 : 9296 - 9306 .
Chang A X , Funkhouser T , Guibas L , et al . ShapeNet: an information-rich 3D model repository [PP/OL ] . V1. arXiv ( 2015-12-09 )[ 2026-04-22 ] . https://arxiv.org/abs/1512.03012 https://arxiv.org/abs/1512.03012 .
Guo Menghao , Cai Junxiong , Liu Zhengning , et al . PCT: point cloud transformer [J ] . Computational Visual Media , 2021 , 7 ( 2 ): 187 - 199 .
Wang Haiyang , Shi Chen , Shi Shaoshuai , et al . DSVT: dynamic sparse voxel transformer with rotated sets [C ] // 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2023 : 13520 - 13529 .
Zhou Wei , Zhang Xiaodan , Hao Xingxing , et al . Multi point-voxel convolution (MPVConv) for deep learning on point clouds [J ] . Computers & Graphics , 2023 , 112 : 72 - 80 .
Liu Zhijian , Tang Haotian , Lin Yujun , et al . Point-voxel CNN for efficient 3D deep learning [C ] // Proceedings of the 33rd International Conference on Neural Information Processing Systems . Red Hook, NY, USA : Curran Associates Inc. , 2019 : 965 - 975 .
Zheng Xiao , Huang Xiaoshui , Mei Guofeng , et al . Point cloud pre-training with diffusion models [C ] // 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2024 : 22935 - 22945 .
Qu Wentao , Shao Yuantian , Meng Lingwu , et al . A conditional denoising diffusion probabilistic model for point cloud upsampling [C ] // 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2024 : 20786 - 20795 .
Liu Yanzhe , Chen Rong , Li Yushi , et al . SPU-PMD: self-supervised point cloud upsampling via progressive mesh deformation [C ] // 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2024 : 5188 - 5197 .
Wang Xuzhi , Feng Wei , Kong Lingdong , et al . NUC-Net: non-uniform cylindrical partition network for efficient LiDAR semantic segmentation [J ] . IEEE Transactions on Circuits and Systems for Video Technology , 2025 , 35 ( 9 ): 9090 - 9104 .
Zhang Jianhui , Luo Yizhi , Zhang Zicheng , et al . CamPoint: boosting point cloud segmentation with virtual camera [C ] // 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2025 : 11822 - 11832 .
Dai Wenxia , Wu Wanmin , Hao Yuqi , et al . A topography-attentional network for ground filtering from ALS point cloud in forest environments [J ] . IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , 2026 , 19 : 1842 - 1855 .
鲁斌 , 尚国栋 . 基于多尺度感知和双模态融合的激光雷达点云语义分割 [J ] . 计算机应用研究 , 2026 , 43 ( 5 ): 1322 - 1328 .
Lu Bin , Shang Guodong . Multi-scale and dual-modality fusion for LiDAR point cloud semantic segmentation [J ] . Application Research of Computers , 2026 , 43 ( 5 ): 1322 - 1328 .
杨军 , 王连甲 . 结合位置关系卷积与深度残差网络的三维点云识别与分割 [J ] . 西安交通大学学报 , 2023 , 57 ( 5 ): 182 - 193 .
Yang Jun , Wang Lianjia . Recognition and segmentation of 3D point cloud through positional relation convolution in combination with deep residual network [J ] . Journal of Xi'an Jiaotong University , 2023 , 57 ( 5 ): 182 - 193 .
Du Zijin , Liang Jianqing , Liang Jiye , et al . Graph regulation network for point cloud segmentation [J ] . IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024 , 46 ( 12 ): 7940 - 7955 .
Qian Guocheng , Li Yuchen , Peng Houwen , et al . PointNeXt: revisiting PointNet++ with improved training and scaling strategies [C ] // Proceedings of the 36th International Conference on Neural Information Processing Systems . Red Hook, NY, USA : Curran Associates Inc. , 2022 : 23192 - 23204 .
Guang Jinzheng , Wu Shichao , Wang Yongru , et al . HTMNet: a hybrid transformer-mamba network for LiDAR-based 3D detection and semantic segmentation [J ] . Expert Systems with Applications , 2026 , 316 : 131832 .
De Vries M , Naidoo R , Fourkioti O , et al . Interpretable point cloud classification using multiple instance learning [C ] // 2025 IEEE/CVF International Conference on Computer Vision (ICCV) . Piscataway, NJ, USA : IEEE , 2025 : 22209 - 22220 .
Wang Zongshun , Li Ce , Feng Zhiqiang , et al . Rethinking static weights: language-guided adaptive weight adjustment for 3D visual grounding [J ] . Knowledge-Based Systems , 2026 , 338 : 115467 .
Zhao Hengshuang , Jiang Li , Jia Jiaya , et al . Point transformer [C ] // 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . Piscataway, NJ, USA : IEEE , 2021 : 16239 - 16248 .
Zhang Cheng , Wan Haocheng , Shen Xinyi , et al . PatchFormer: an efficient point transformer with patch attention [C ] // 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2022 : 11789 - 11798 .
唐友源 , 张辉 , 杜瑞 , 等 . 基于结构频谱感知框架的配电网点云语义分割 [J ] . 自动化学报 , 2026 , 52 ( 4 ): 833 - 845 .
Tang Youyuan , Zhang Hui , Du Rui , et al . Semantic segmentation of distribution network point clouds based on a structure spectrum-aware framework [J ] . Acta Automatica Sinica , 2026 , 52 ( 4 ): 833 - 845 .
Zhang Weijian , Song Haichuan , Zhang Zhizhong , et al . from sparse semantics to rich instances: empowering label-efficient LiDAR panoptic segmentation via geometric priors [J ] . Neural Networks , 2026 , 200 : 108767 .
朱安迪 , 达飞鹏 , 盖绍彦 . 对融合特征敏感的三维点云识别与分割 [J ] . 西安交通大学学报 , 2024 , 58 ( 5 ): 52 - 63 .
Zhu Andi , Da Feipeng , Ge Shaoyan . Recognition and segmentation of 3D point clouds sensitive to fusion features [J ] . Journal of Xi'an Jiaotong University , 2024 , 58 ( 5 ): 52 - 63 .
Zhang Cheng , Wan Haocheng , Liu Shengqiang , et al . Point-voxel transformer: an efficient approach to 3D deep learning [PP/OL ] . V1. arXiv ( 2021-08-13 )[ 2026-04-22 ] . https://arxiv.org/abs/2108.06076v1 https://arxiv.org/abs/2108.06076v1 .
Shi Shaoshuai , Guo Chaoxu , Jiang Li , et al . PV-RCNN: point-voxel feature set abstraction for 3D object detection [C ] // 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2020 : 10526 - 10535 .
Shi Shaoshuai , Jiang Li , Deng Jiajun , et al . PV-RCNN++: point-voxel feature set abstraction with local vector representation for 3D object detection [J ] . International Journal of Computer Vision , 2023 , 131 ( 2 ): 531 - 551 .
Qi C R , Yi Li , Su Hao , et al . PointNet++: deep hierarchical feature learning on point sets in a metric space [C ] // Proceedings of the 31st International Conference on Neural Information Processing Systems . Red Hook, NY, USA : Curran Associates Inc. , 2017 : 5105 - 5114 .
Milioto A , Vizzo I , Behley J , et al . RangeNet ++: fast and accurate LiDAR semantic segmentation [C ] // 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . Piscataway, NJ, USA : IEEE , 2019 : 4213 - 4220 .
Hu Qingyong , Yang Bo , Xie Linhai , et al . RandLA-Net: efficient semantic segmentation of large-scale point clouds [C ] // 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2020 : 11105 - 11114 .
Xu Chenfeng , Wu Bichen , Wang Zining , et al . SqueezeSegV3: spatially-adaptive convolution for efficient Point-Cloud segmentation [C ] // Computer Vision – ECCV 2020 . Cham : Springer International Publishing , 2020 : 1 - 19 .
Thomas H , Qi C R , Deschaud J E , et al . KPConv: flexible and deformable convolution for point clouds [C ] // 2019 IEEE/CVF International Conference on Computer Vision (ICCV) . Piscataway, NJ, USA : IEEE , 2019 : 6410 - 6419 .
Cortinhal T , Tzelepis G , Erdal Aksoy E . SalsaNext: fast, uncertainty-aware semantic segmentation of LiDAR point clouds [C ] // Advances in Visual Computing . Cham : Springer International Publishing , 2020 : 207 - 222 .
Ye Maosheng , Xu Shuangjie , Cao Tongyi , et al . DRINet: a dual-representation iterative learning network for point cloud segmentation [C ] // 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . Piscataway, NJ, USA : IEEE , 2021 : 7427 - 7436 .
Zhu Xinge , Zhou Hui , Wang Tai , et al . Cylindrical and asymmetrical 3D convolution networks for LiDAR segmentation [C ] // 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2021 : 9934 - 9943 .
Ye Maosheng , Wan Rui , Xu Shuangjie , et al . Efficient point cloud segmentation with geometry-aware sparse networks [C ] // Computer Vision – ECCV 2022 . Cham : Springer Nature Switzerland , 2022 : 196 - 212 .
Xu Jianyun , Zhang Ruixiang , Dou Jian , et al . RPVNet: a deep and efficient range-point-voxel fusion network for LiDAR point cloud segmentation [C ] // 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . Piscataway, NJ, USA : IEEE , 2021 : 16004 - 16013 .
Li Yangyan , Bu Rui , Sun Mingchao , et al . PointCNN: convolution on Χ -transformed points [C ] // Proceedings of the 32nd International Conference on Neural Information Processing Systems . Red Hook, NY, USA : Curran Associates Inc. , 2018 : 828 - 838 .
Yan Xu , Zheng Chaoda , Li Zhen , et al . PointASNL: robust point clouds processing using nonlocal neural networks with adaptive sampling [C ] // 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2020 : 5588 - 5597 .
Ma Xu , Qin Can , You Haoxuan , et al . Rethinking network design and local geometry in point cloud: a simple residual MLP framework [C ] // ICLR 2022 . New York, USA : ICLR , 2022 : 1 - 15 .
Xu Mutian , Ding Runyu , Zhao Hengshuang , et al . PAConv: position adaptive convolution with dynamic kernel assembling on point clouds [C ] // 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway, NJ, USA : IEEE , 2021 : 3172 - 3181 .
Zheng Qiang , Zhang Chao , Sun Jian . PointMT: efficient point cloud analysis with hybrid MLP-transformer architecture [J ] . IEEE Transactions on Multimedia , 2025 , 27 : 6382 - 6396 .
0
浏览量
0
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621