高级检索

文本引导的多模态融合类别级物体位姿估计方法

A Text-guided Multi-modal Fusion Method for Category-level Object Pose Estimation

  • 摘要: 类别级6D物体位姿估计能够泛化到同一类别内未见过的实例,在机器人操作、增强现实和场景理解等应用中受到广泛关注。针对传统类别级位姿估计方法多依赖RGB图像和点云等视觉模态信息,在纹理缺失、光照变化或遮挡等复杂场景下鲁棒性不足,且未充分利用物体语义先验知识的问题,提出一种文本引导的多模态融合类别级物体位姿估计方法(A Text-Guided Multi-Modal Fusion Method for Category-Level Object Pose Estimation,TMF-Pose)。首先将文本语义信息融入位姿估计流程,通过预训练CLIP模型提取物体类别结构化描述的高维语义特征,设计文本引导的多模态特征融合模块,并结合双线性门控机制与交叉注意力机制,实现RGB特征、点云几何特征与文本语义特征的自适应对齐与互补融合;然后构建实例自适应特征增强关键点检测模块,并通过几何感知特征聚合模块整合局部与全局几何信息;最后引入形状重建模块生成密集归一化形状表示,以几何一致性约束优化位姿估计结果。实验结果表明,在REAL275、CAMERA25和HouseCat6D数据集上,所提方法的5°2 cm指标分别达到61.4%、78.8%和22.6%,较各数据集的最优对比结果分别提高0.6、0.8和1.2个百分点,验证了所提方法在类别级物体位姿估计任务中的有效性。

     

    Abstract: Category-level 6D object pose estimation plays an important role in robotics, augmented reality, and scene understanding because it enables generalization to unseen instances within the same category. Existing category-level pose estimation methods mainly rely on RGB images and point clouds, which often lack ro-bustness under challenging conditions such as texture-less objects, illumination variations, and occlusions. Moreover, object semantic priors are not fully exploited. To address these issues, TMF-Pose, a text-guided multi-modal fusion method for category-level object pose estimation, is proposed. The proposed method first incorporates textual semantic information into the pose estimation pipeline and extracts high-dimensional semantic features from structured object category descriptions using a pre-trained CLIP model. The text-guided fusion module combines bilinear gating with cross-attention. It aligns and inte-grates RGB, point cloud, and text features. Furthermore, an instance-adaptive keypoint detection module with feature enhancement is constructed, and local and global geometric information are integrated through a geometry-aware feature aggregation module. Finally, a shape reconstruction module is intro-duced to generate dense normalized shape representations and optimize pose estimation results through geometric consistency constraints. On REAL275, CAMERA25, and HouseCat6D, TMF-Pose achieved 5°2 cm accuracies of 61.4%, 78.8%, and 22.6%, respectively. These results exceeded the best baseline result on each dataset by 0.6, 0.8, and 1.2 percentage points, respectively.

     

/

返回文章
返回