Abstract:
Category-level 6D object pose estimation plays an important role in robotics, augmented reality, and scene understanding because it enables generalization to unseen instances within the same category. Existing category-level pose estimation methods mainly rely on RGB images and point clouds, which often lack ro-bustness under challenging conditions such as texture-less objects, illumination variations, and occlusions. Moreover, object semantic priors are not fully exploited. To address these issues, TMF-Pose, a text-guided multi-modal fusion method for category-level object pose estimation, is proposed. The proposed method first incorporates textual semantic information into the pose estimation pipeline and extracts high-dimensional semantic features from structured object category descriptions using a pre-trained CLIP model. The text-guided fusion module combines bilinear gating with cross-attention. It aligns and inte-grates RGB, point cloud, and text features. Furthermore, an instance-adaptive keypoint detection module with feature enhancement is constructed, and local and global geometric information are integrated through a geometry-aware feature aggregation module. Finally, a shape reconstruction module is intro-duced to generate dense normalized shape representations and optimize pose estimation results through geometric consistency constraints. On REAL275, CAMERA25, and HouseCat6D, TMF-Pose achieved 5°2 cm accuracies of 61.4%, 78.8%, and 22.6%, respectively. These results exceeded the best baseline result on each dataset by 0.6, 0.8, and 1.2 percentage points, respectively.