吉林大学学报(工学版) ›› 2026, Vol. 56 ›› Issue (2): 523-532.doi: 10.13229/j.cnki.jdxbgxb.20240748

• 计算机科学与技术 • 上一篇    

基于双向特征融合的物体位姿估计方法

缪君1,2(),严杰1,杜荣华1(),李磊3,储珺3   

  1. 1.南昌航空大学 航空制造与机械工程学院,南昌 330063
    2.中国科学院月球与深空探测重点实验室,北京 100190
    3.南昌航空大学 计算机视觉研究所,南昌 330063
  • 收稿日期:2024-07-06 出版日期:2026-02-01 发布日期:2026-03-17
  • 通讯作者: 杜荣华 E-mail:miaojun@nchu.edu.cn;7130l@nchu.edu.cn
  • 作者简介:缪君(1979-),男,教授,博士.研究方向:三维场景理解与重建,工业视觉检测.E-mail:miaojun@nchu.edu.cn
  • 基金资助:
    国家自然科学基金项目(62366032);国家自然科学基金项目(62361043);国家自然科学基金项目(62162045);中国科学院月球与深空探测重点实验室开放基金项目(LDSE202301)

A bidirectional feature fusion method for object position estimation

Jun MIAO1,2(),Jie YAN1,Rong-hua DU1(),Lei LI3,Jun CHU3   

  1. 1.School of Aeronautical Manufacturing and Mechanical Engineering,Nanchang Hangkong University,Nanchang 330063,China
    2.Key Laboratory of Lunar and Deep Space Exploration,CAS,Beijing 100190,China
    3.Institute of Computer Vision,Nanchang Hangkong University,Nanchang 330063,China
  • Received:2024-07-06 Online:2026-02-01 Published:2026-03-17
  • Contact: Rong-hua DU E-mail:miaojun@nchu.edu.cn;7130l@nchu.edu.cn

摘要:

为充分利用RGB图像的外观特征和深度图像的几何特征,提出了一种“外观-几何”特征并行融合的物体位姿估计方法。首先,在特征提取与融合阶段,构建了一种具有3个并行支流的双向融合体系结构,确保在每个编码层和解码层对并行的RGB图像特征和深度图像特征进行融合;同时,为避免重要特征丢失,且实现两种特征的充分融合,设计了两个互补的注意力机制,使两种特征获得局部和全局的互补;其次,在位姿推理计算阶段,考虑网络输出关键点与物体中心点之间的距离,提出了一种基于距离量和距离约束相结合的关键点检测网络,实现了精确的位姿估计。本文算法其在两个具有挑战性的6D物体位姿估计数据集上进行了测试,验证了其有效性。

关键词: 位姿估计, 双向特征融合, 特征差异, 距离约束, 注意力机制

Abstract:

To fully leverage the appearance features of RGB images and the geometric features of depth images, this paper proposes an "appearance-geometry" features parallel fusion method for object position estimation. First, in the feature extraction and fusion stage, a three-stream bidirectional fusion architecture is constructed to ensure that the parallel RGB image features and depth image features are fused at each encoding layer and decoding layer. To prevent the loss of important features and achieve sufficient fusion of the two types of features, two complementary attention mechanisms are designed, enabling the two features to gain both local and global complementarities. Sercond, in the pose inference calculation stage, considering the distance between the keypoints output by the network and the object’s center point, a keypoint detection network based on a combination of distance metric and distance constraint is proposed, achieving accurate position estimation. The proposed algorithm has been tested on two challenging 6D object position estimation datasets, validating its effectiveness.

Key words: position estimation, bidirectional feature fusion, feature disparity, distance constraint, attention mechanism

中图分类号: 

  • TP391.41

图1

本文提出的位姿估计网络结构图"

图2

本文提出的特征处理模块结构图"

图3

本文提出的关键点检测网络(DCKP)结构图"

表1

在LineMOD数据集上,本文方法与其他方法对比 (%)"

RGB-D
DenseFusionPVN3DFFB6D本文方法(KP)本文方法
objectADDADDSADDADDSADDADDSADDADDSADDADDS
ape83.970.290.984.696.080.598.490.699.896.5
benchvise86.983.193.189.396.194.898.394.698.595.1
camera84.288.494.394.297.496.398.696.699.698.6
can86.076.290.584.898.288.597.689.697.996.8
cat90.481.095.687.597.596.299.897.099.897.6
driller88.070.791.783.896.089.398.893.999.697.0
duck84.182.794.387.197.195.799.194.699.397.8
eggbox87.285.292.989.597.796.198.196.999.499.4
gule83.574.588.281.693.388.698.794.199.799.7
holepuncher86.082.394.881.096.693.799.295.599.595.3
iron87.083.689.587.397.496.599.696.999.297.3
lamp81.675.388.479.896.093.298.896.399.597.7
phone79.679.682.388.386.090.296.396.399.698.6
ALL85.379.491.386.195.892.398.694.899.497.5

图4

在LineMOD数据集上,本文方法与其他方法可视化结果"

表2

本文所提关键点检测网络的有效性验证 (%)"

方法Baseline+ICPBaseline+KPBaseline+DCKP
ADD93.095.597.4
ADDS86.991.894.5

表3

在YCB-Video数据集上,本文方法与其他方法对比 (%)"

RGBRGB-D
PoseCNNPVNetDenseFusionFFB6D本文方法
objectADDADDSADDADDSADDADDSADDADDSADDADDS
002_master_chef_can83.950.290.974.695.370.796.380.698.293.2
003_cracker_box76.953.187.179.392.586.996.394.699.896.3
004_sugar_box84.268.494.384.295.190.897.696.699.297.9
005_tomato_soup_can81.066.290.579.893.884.795.689.697.894.8
006_mustard_bottle90.481.090.683.595.890.997.897.098.898.5
007_tuna_fish_can88.070.791.773.895.779.696.889.799.888.9
008_pudding_box79.162.789.384.194.389.397.194.697.796.8
009_gelatin_box87.275.292.989.597.295.898.396.998.197.3
024_bowl*69.669.680.380.386.086.096.396.399.899.8
025_mug78.258.590.776.695.383.897.394.298.593.4
035_power_drill72.755.387.478.492.183.797.295.998.097.6
036_wood_block*64.364.384.284.289.589.592.692.698.998.9
037_scissors56.935.884.270.390.177.497.795.798.595.2
040_large_maker*71.758.389.581.095.189.196.689.199.099.0
051_large_clamp*50.250.263.663.671.571.596.896.899.199.1
061_foam_brick88.088.083.183.192.292.297.397.398.698.6
ALL76.463.086.979.192.085.196.793.698.896.6

图5

在YCB-Video数据集上,本文方法训练的可视化结果"

[1] Guan J, Hao Y M, Wu Q X, et al. A survey of 6DoF object pose estimation methods for different application scenarios[J]. Sensors, 2024, 24(4): 1076.
[2] Marullo G, Tanzi L, Piazzolla P, et al. 6D object position estimation from 2D images: A literature review[J]. Multimedia Tools and Applications, 2023, 82(16): 24605-24643.
[3] 王静, 金玉楚, 郭苹, 等. 基于深度学习的相机位姿估计方法综述[J]. 计算机工程与应用, 2023, 59(7): 1-14.
Wang Jing, Jin Yu-chu, Guo Ping, et al. A review of camera pose estimation methods based on deep learning[J]. Computer Engineering and Applications, 2023, 59(7): 1-14.
[4] Wang C, Xu D E, Zhu Y K, et al. Dense Fusion: 6D object pose estimation by iterative dense fusion[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE,2019: 3343-3352.
[5] He Y S, Huang H B, Fan H Q, et al. FB6D: A full flow bidirectional fusion network for 6D pose estimation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEEE, 2021: 3003-3013.
[6] Peng S D, Liu Y, Huang Q X, et al. PVNet: Pixel-wise voting network for 6DoF pose estimation[J]. Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(6): 3212-3223.
[7] Lin S F, Wang Z R, Ling Y G, et al. E2EK: End-to-end regression network based on keypoint for 6D pose estimation[J]. IEEE Robotics and Automation Letters, 2022, 7(3): 6526-6533.
[8] Xiang Y, Schmidt T, Narayanan V, et al. Pose CNN: A convolutional neural network for 6D object pose estimation in cluttered scenes[J]. ArXiv Preprint, 2017, 11: 171100199.
[9] Zakharov S, Shugurov I, Ilic S. DPOD: 6D pose object detector and refiner[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway, NJ: IEEE, 2019: 1941-1950.
[10] 王连明, 吴鑫. 基于姿态估计的物体 3D 运动参数测量方法[J]. 吉林大学学报:工学版, 2023, 53(7): 2099-2108.
Wang Lian-ming, Wu Xin. Measurement of 3D motion parameters of an object based on attitude estimation[J]. Journal of Jilin University (Engineering and Technology Edition), 2023, 53(7): 2099-2108.
[11] Ding Z F, Sun Y X, Xu S J, et al. Recent advances and perspectives in deep learning techniques for 3D point cloud data processing[J]. Robotics, 2023, 12(4): 100.
[12] Zhou J, Chen K, Xu L L, et al. Deep fusion transformer network with weighted vector-wise keypoints voting for robust 6D object pose estimation[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway, NJ: IEEE, 2023: 13967-13977.
[13] 白琳, 刘林军, 李轩昂, 等. 基于自监督学习的单目图像深度估计算法[J]. 吉林大学学报:工学版, 2023, 53(4): 1139-1145.
Bai Lin, Liu Lin-jun, Li Xuan-ang, et al. A depth estimation algorithm for monocular images based on self-supervised learning[J]. Journal of Jilin University (Engineering and Technology Edition), 2023, 53(4): 1139-1145.
[14] Song C, Song J R, Huang Q X. HybridPose: 6D object pose estimation under hybrid representations[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2020: 431-440.
[15] 张宸嘉, 朱磊, 俞璐. 卷积神经网络中的注意力机制综述[J]. 计算机工程与应用学报, 2021, 57(20):64-72.
Zhang Chen-jia, Zhu Lei, Yu Lu. A review of attention mechanisms in convolutional neural networks[J]. Journal of Computer Engineering & Applications, 2021, 57(20):64-72.
[16] Hinterstoisser S, Lepetit V, Ilic S, et al. Model based training, detection and pose estimation of texture-less 3D objects in heavily cluttered scenes[C]∥Computer Vision-ACCV 2012: 11th Asian Conference on Computer Vision. Piscataway, NJ: IEEE, 2013: 548-562.
[17] Calli B, Singh A, Walsman A, et al. The YCB object and model set: Towards common benchmarks for manipulation research[C]∥ International Conference on Advanced Robotics. Piscataway, NJ: IEEE, 2015: 510-517.
[1] 周求湛,李新萌,沈皓庆子,武慧南,李媛媛,荣静,胡春华,刘萍萍. 基于注意力机制的不平衡数据的非侵入式负荷分解[J]. 吉林大学学报(工学版), 2026, 56(1): 239-246.
[2] 冯志刚,任梦媛,董冰,于明月. 基于多频带特征图和改进SqueezeNet的滚动轴承故障诊断[J]. 吉林大学学报(工学版), 2026, 56(1): 96-108.
[3] 霍震,金立生,华强,贺阳. 基于边缘特征引导的智能汽车语义分割方法[J]. 吉林大学学报(工学版), 2025, 55(9): 3032-3041.
[4] 庄珊娜,王君帅,白晶,杜京瑾,王正友. 基于三维卷积与自注意力机制的视频行人重识别[J]. 吉林大学学报(工学版), 2025, 55(7): 2409-2417.
[5] 冯志刚,王首起,于明月. 基于变分模态提取及轻量级网络的滚动轴承故障诊断[J]. 吉林大学学报(工学版), 2025, 55(6): 1883-1891.
[6] 薛雅丽,俞潼安,崔闪,周李尊. 基于级联嵌套U-Net的红外小目标检测[J]. 吉林大学学报(工学版), 2025, 55(5): 1714-1721.
[7] 才华,王玉瑶,付强,马智勇,王伟刚,张晨洁. 基于注意力机制和特征融合的语义分割网络[J]. 吉林大学学报(工学版), 2025, 55(4): 1384-1395.
[8] 张河山,范梦伟,谭鑫,郑展骥,寇立明,徐进. 基于改进YOLOX的无人机航拍图像密集小目标车辆检测[J]. 吉林大学学报(工学版), 2025, 55(4): 1307-1318.
[9] 张兰芳,李根泽,刘婷宇,余博. 局部多车影响下跟驰行为机理及建模[J]. 吉林大学学报(工学版), 2025, 55(3): 963-973.
[10] 李扬,李现国,苗长云,徐晟. 基于双分支通道先验和Retinex的低照度图像增强算法[J]. 吉林大学学报(工学版), 2025, 55(3): 1028-1036.
[11] 王祥,谭国真,彭衍飞,任浩,李健平. 基于语言推理和认知记忆的自动驾驶决策模型[J]. 吉林大学学报(工学版), 2025, 55(12): 3918-3927.
[12] 李云红,王梅,苏雪平,李丽敏,张富星,郝特吉. 结合注意力与上下文融合的遥感图像道路提取[J]. 吉林大学学报(工学版), 2025, 55(12): 4034-4044.
[13] 彭铎,刘明硕,谢堃. 观测站参数误差下联合多特征融合注意力机制的TDOA/FDOA多机无源定位算法[J]. 吉林大学学报(工学版), 2025, 55(11): 3751-3761.
[14] 周求湛,牟岩,武慧南,陈霄,汪锋,李琛,张雯,刘萍萍,王聪. 基于RBVS和CBCNN的风机叶片故障检测和分类方法[J]. 吉林大学学报(工学版), 2025, 55(10): 3119-3130.
[15] 郭晓然,王铁君,闫悦. 基于局部注意力和本地远程监督的实体关系抽取方法[J]. 吉林大学学报(工学版), 2025, 55(1): 307-315.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!