Abstract:
To address the insufficient discriminability of image features, the limited semantic expressiveness of point cloud features, and the low cross-modal correspondence accuracy in existing methods for correspondence calculation from 2D images to 3D point clouds under low-overlap and high-noise scenarios, a dense correspondence calculation method based on a dual-modal feature optimization mechanism is proposed in this paper. First, a feature discriminability optimization module is designed, which combines normalization-guided feature evaluation mechanism and channel reconstruction mechanism to dynamically select and enhance image features, so as to alleviate the weakening of target features caused by background interference in high-noise scenes. Second, a symmetric overlap region detection architecture is adopted to predict the effective overlapping regions between images and point clouds, thereby reducing the matching search range in low-overlap scenes. Then, a cross-modal feature correspondence module is designed, which combines a local geometric context fusion mechanism and a global bilinear regularization mechanism to reinforce the feature expressiveness of 3D point clouds at the structural and semantic levels. Finally, a nearest-neighbor strategy in feature space is employed to establish robust cross-modal dense correspondences between images and point clouds within their spatially overlapping regions. Experiments on the KITTI and NuScenes datasets demonstrate that our method achieves Relative Translational Errors (RTE) of 0.85 m and 1.68 m, and Relative Rotational Errors (RRE) of 2.06° and 2.70°, respectively. Comparative results with current mainstream methods show that the proposed approach significantly improves the accuracy of cross-modal dense correspondence calculation from 2D images to 3D point clouds and exhibits strong generalization capability.