北京邮电大学智能工程与自动化学院,100876,北京
收稿:2026-03-24,
修回:2026-07-06,
录用:2026-07-07,
移动端阅览
张怡阳, 虞承舜, 李彤, 等. 一种视触觉融合驱动的抓取目标形状重建方法[J/OL]. 西安交通大学学报, 2026.
ZHANG Yiyhang, YU Chengshun, LI Tong, et al. A Method of Reconstructing the Shape of Grasping Target Driven by Visual-Tactile Fusion[J/OL]. JOURNAL OF XI’AN JIAOTONG UNIVERSITY, 2026.
针对非结构化家庭服务场景下,机器人抓取目标形状重建面临的抓取目标特性多样化和视线遮挡等挑战,提出一种视触觉融合驱动的抓取目标形状重建方法。首先,设计块特征转化网络,将单视角输入视觉点云与触觉阵列转化为块特征代理。其次,利用Transformer架构实现视触觉跨模态特征融合与补全,设计几何感知模块预测模拟局部几何关系。最后,结合形状查询器生成基于动态条件的查询嵌入,设计解码器折叠重建模块,补全基于块特征预测的点云,输出完整目标形状点云。实验结果表明:所提出方法对机器人执行抓取任务的抓取成功率有显著提升,形状重建性能相较于现有最佳基线模型提升了12.5%,在陌生物体上的平均准确度高达80.3%,同时触觉模态的融入使模型的倒角距离下降了1.532。所提方法通过跨模态数据互补可以实现更完整、鲁棒性更强的抓取目标点云形状重建,且在处理复杂多变目标时具有较高的鲁棒性与泛化能力,为非结构化家庭服务场景下,机器人抓取目标的形状重建提供了一种解决方案。
To address the challenges of diverse target characteristics and visual occlusion in grasping target shape reconstruction for robots in unstructured home service scenarios
this paper proposes a shape reconstruction method for grasping targets driven by visual-tactile fusion. First
a block feature transformation network is designed to transform the single-view input point cloud and the tactile array into patch feature proxies. Second
the Transformer architecture is utilized to achieve visual-tactile cross-modal feature fusion and completion
and a geometric perception module is designed to predict and model local geometric relationships. Finally
combined with a shape query module to generate dynamic condition-based query embeddings
a decoder folding reconstruction module is designed to complete the point cloud based on patch feature predictions
thereby outputting the complete target shape point cloud. The experimental results demonstrate that the proposed method significantly improves the grasp success rate of robotic grasping tasks. Compared with the current best baseline model
the proposed approach achieves a 12.5% improvement in shape reconstruction performance and attains an average accuracy of up to 80.3% on unseen objects. Furthermore
the integration of the tactile modality reduces the Chamfer distance by 1.532. Through cross-modal data complementarity
the proposed method achieves more complete and robust point cloud shape reconstruction for grasping targets
demonstrating strong robustness and generalization capabilities when handling complex and variable objects
thereby providing a solution for shape reconstruction of grasping targets for robots in unstructured home service scenarios.
LIU L , Yang W , Fei B . GeeNet: robust and fast point cloud completion for ground elevation estimation towards autonomous vehicles [J ] . Frontiers of Information Technology & Electronic Engineering , 2024 , 25 ( 7 ): 938 - 950 .
武湘怡 , 叶海良 , 曹飞龙 . 基于平滑-锐化图卷积的点云补全方法 [J/OL ] . 计算机应用 , 1 - 12 [ 2026-03-20 ] . https://link.cnki.net/urlid/51.1307.TP.20250924.1815.006 https://link.cnki.net/urlid/51.1307.TP.20250924.1815.006 .
WU Xiangyi , YE Hailiang , CAO Feilong . Point cloud completion method based onsmooth-sharpen graph convolution [J/OL ] . Journal of Computer Applications , 1 - 12 [ 2026-03-20 ] . https://link.cnki.net/urlid/51.1307.TP.20250924.1815.006 https://link.cnki.net/urlid/51.1307.TP.20250924.1815.006 .
陆春媚 , 杨志景 . 多级精细化反卷积点云补全网络 [J ] . Journal of Computer Engineering & Applications , 2023 , 59 ( 17 ).
LU Chunmei , YANG Zhijing . Multistage Refinement of Deconvolution Point Cloud Complementation Network [J ] . Journal of Computer Engineering & Applications , 2023 , 59 ( 17 ).
PISTILLI F , Fracastoro G , Valsesia D , et al . Learning robust graph-convolutional representations for point cloud denoising [J ] . IEEE Journal of Selected Topics in Signal Processing , 2020 , 15 ( 2 ): 402 - 414 .
石键瀚 . 基于融合触觉和视觉的点云补全理论研究 [D ] . 南京邮电大学 , 2024 . DOI: 10.27251/d.cnki.gnjdc.2024.000420 http://dx.doi.org/10.27251/d.cnki.gnjdc.2024.000420 .
陈仁祥 , 邱天然 , 杨黎霞 , 等 . 基于空间信息聚合的遮挡目标抓取位姿检测 [J ] . Optics and Precision Engineering , 2024 , 32 ( 18 ): 2792 - 2802 .
CHEN Renxiang , QIU Tianran , YANG Lixia . Occludedtarget graspingdetection method based on spatial information aggregation [J ] . Optics and Precision Engineering , 2024 , 32 ( 18 ): 2792 - 2802 .
SHENG Q , Zhou Z , Li J , et al . A Comprehensive Review of Humanoid Robots [J ] . SmartBot , 2025 , 1 ( 1 ): e12008 .
Li T , Yan Y , Yu C , et al . A comprehensive review of robot intelligent grasping based on tactile perception [J ] . Robotics and Computer-Integrated Manufacturing , 2024 , 90 : 102792 .
BAUZA M , Canal O , Rodriguez A . Tactile maping and localization from high-resolution tactile imprints [C ] // 2019 International Conference on Robotics and Automation (ICRA) . IEEE , 2019 : 3811 - 3817 .
DONLON E , DONG S Y , LIU M , et al . GelSlim: a high-resolution, compact, robust, and calibrated tactile-sensing finger [C ] // Proceedings of 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems . Piscataway : IEEE Press , 2018 : 1927 - 1934 .
SURESH S , Si Z , Mangelson J G , et al . ShapeMap 3-D: Efficient shape map** through dense touch and vision [C ] // 2022 International Conference on Robotics and Automation (ICRA) . IEEE , 2022 : 7073 - 7080 ..
OTTENHAUS S , Miller M , Schiebener D , et al . Local implicit surface estimation for haptic exploration [C ] // 2016 IEEE-RAS 16th International Conference on Humanoid Robots (Humanoids) . IEEE , 2016 : 850 - 856 .
ZHANG P , Bai L , Shan D , et al . Visual–tactile fusion object classification method based on adaptive feature weighting [J ] . International Journal of Advanced Robotic Systems , 2023 , 20 ( 4 ): 17298806231191947 .
SMITH E , Meger D , Pineda L , et al . Active 3D shape reconstruction from vision and touch [J ] . Advances in Neural Information Processing Systems , 2021 , 34 : 16064 - 16078 .
WATKINS-VALLS D , Varley J , Allen P . Multi-modal geometric learning for grasping and manipulation [C ] // 2019 International conference on robotics and automation (ICRA) . IEEE , 2019 : 7339 - 7345 .
OTTENHAUS S , Renninghoff D , Grimm R , et al . Visuo-haptic grasping of unknown objects based on gaussian process implicit surfaces and deep learning [C ] // 2019 IEEE-RAS 19th International Conference on Humanoid Robots (Humanoids) . IEEE , 2019 : 402 - 409 .
CHANG A X , Funkhouser T , Guibas L , et al . Shapenet: An information-rich 3d model repository [J ] . arXiv preprint arXiv: 1512.03012 , 2015 .
LI T , Yan Y , Yu C , et al . VTG: A Visual-Tactile Dataset for Three-Finger Grasp [J ] . IEEE Robotics and Automation Letters , 2024 .
Fang H S , Wang C , Gou M , et al . Graspnet-1billion: A large-scale benchmark for general object grasping [C ] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2020 : 11444 - 11453 .
YU X , Rao Y , Wang Z , et al . Pointr: Diverse point cloud completion with geometry-aware transformers [C ] // Proceedings of the IEEE/CVF international conference on computer vision . 2021 : 12498 - 12507 .
XIANG P , Wen X , Liu Y S , et al . Snowflakenet: Pointcloud completion by snowflake point deconvolution with skip-transformer [C ] // Proceedings of the IEEE/CVF international conference on computer vision . 2021 : 5499 - 5509 .
TANG J , Gong Z , Yi R , et al . Lake-net: Topology-aware point cloud completion by localizing aligned keypoints [C ] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2022 : 1726 - 1735 .
ZHOU H , Cao Y , Chu W , et al . Seedformer: Patch seeds based point cloud completion with upsample transformer [C ] // European conference on computer vision . Cham : Springer Nature Switzerland , 2022 : 416 - 432 .
YU X , Rao Y , Wang Z , et al . AdaPoinTr: Diverse point cloud completion with adaptive geometry-aware transformers [J ] . IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023 , 45 ( 12 ): 14114 - 14130 .
Li J , Guo S , Wang L , et al . CompleteDT: Point cloud completion with information-perception transformers [J ] . Neurocomputing , 2024 , 592 : 127790 .
Li Y , Zhou Q , Gong J , et al . Dapointr: Domain adaptive point transformer for point cloud completion [C ] // Proceedings of the AAAI Conference on Artificial Intelligence . 2025 , 39 ( 5 ): 5066 - 5074 .
Yu X , Li J , Wong C C , et al . FACNet: Feature alignment fast point cloud completion network [J ] . Computational Visual Media , 2025 , 11 ( 1 ): 141 - 157 .
0
浏览量
1
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621