高级检索

基于双模态特征优化的二维图像到三维点云密集对应关系计算

Dense Correspondence Calculation from 2D Images to 3D Point Clouds via Dual-Modal Feature Optimization

  • 摘要: 针对现有二维图像到三维点云对应关系计算方法在低重叠、强噪声场景中存在的图像特征判别性不足、点云语义表达能力弱、跨模态对应精度低等问题,提出基于双模态特征优化机制的二维图像到三维点云密集对应关系计算方法。首先,设计特征判别性优化模块,结合归一化引导的特征评估与通道重构机制对图像特征进行动态筛选与增强,缓解强噪声场景下背景干扰导致的目标特征弱化问题;然后,通过对称式重叠区域检测架构预测图像与点云之间的有效重叠区域,缩小低重叠场景下的匹配搜索范围;其次,设计跨模态特征对应模块,结合局部几何上下文融合与全局双线性正则化机制,从结构与语义层面强化三维点云特征表达能力;最后,基于特征空间的最近邻策略,在图像与点云的空间重叠区域中构建鲁棒的跨模态密集对应关系。在KITTI数据集和NuScenes数据集上的实验结果表明,所提方法的相对平移误差分别为0.85 m和1.68 m,相对旋转误差分别为2.06°和2.70°;该方法能有效地提高二维图像到三维点云的跨模态密集对应关系计算准确率,具有良好的泛化能力。

     

    Abstract: To address the insufficient discriminability of image features, the limited semantic expressiveness of point cloud features, and the low cross-modal correspondence accuracy in existing methods for correspondence calculation from 2D images to 3D point clouds under low-overlap and high-noise scenarios, a dense correspondence calculation method based on a dual-modal feature optimization mechanism is proposed in this paper. First, a feature discriminability optimization module is designed, which combines normalization-guided feature evaluation mechanism and channel reconstruction mechanism to dynamically select and enhance image features, so as to alleviate the weakening of target features caused by background interference in high-noise scenes. Second, a symmetric overlap region detection architecture is adopted to predict the effective overlapping regions between images and point clouds, thereby reducing the matching search range in low-overlap scenes. Then, a cross-modal feature correspondence module is designed, which combines a local geometric context fusion mechanism and a global bilinear regularization mechanism to reinforce the feature expressiveness of 3D point clouds at the structural and semantic levels. Finally, a nearest-neighbor strategy in feature space is employed to establish robust cross-modal dense correspondences between images and point clouds within their spatially overlapping regions. Experiments on the KITTI and NuScenes datasets demonstrate that our method achieves Relative Translational Errors (RTE) of 0.85 m and 1.68 m, and Relative Rotational Errors (RRE) of 2.06° and 2.70°, respectively. Comparative results with current mainstream methods show that the proposed approach significantly improves the accuracy of cross-modal dense correspondence calculation from 2D images to 3D point clouds and exhibits strong generalization capability.

     

/

返回文章
返回