高级检索

面向特定对象的个性化文生三维模型

Text-guided Customized 3D Model Generation for Specific-subjects

  • 摘要: 针对特定对象的文本生成三维模型任务,现有方法难以在身份一致性与视角多样性之间取得平衡。为此,提出一种面向特定对象的个性化文生三维模型方法:首先,微调扩散模型,利用特定对象图像学习其身份特征,得到专属的个性化图像生成模型;其次,通过大语言模型分析输入描述,结合高斯初始化模块生成干净、主体突出的正视角图像;再次,将布局控制信息注入扩散模型的交叉注意力层,约束图像中对象的位置与完整性;最后,借助图像生成三维模型技术,由正视角图像预测多视角信息,融合重建最终三维模型。在DreamBooth数据集上的实验结果表明,所提方法的CLIP-I得分为0.697,CLIP-T得分为0.295,在三维模型的结构精度与视觉质量方面均显著优于DreamBooth3D、MVDream、GaussianEditor等现有方法,验证了该方法的有效性。

     

    Abstract: The task of text-guided customized 3D model generation for specific-subjects poses a fundamental chal-lenge: existing methods struggle to balance identity consistency with viewpoint diversity. To address this, a personalized text-to-3D model generation method is proposed for specific subjects. First, a diffusion model is fine-tuned on images of the target subject to learn its identity features, yielding a subject-specific per-sonalized image generation model; second, a large language model analyzes the input text description, and a Gaussian-based initialization module is employed to generate a clean front-view image with a prominent subject; third, layout control information is injected into the cross-attention layers of the diffusion model to constrain the position and completeness of the subject in the generated image; finally, an image-to-3D modeling technique is applied to predict multi-view information from the front-view image and reconstruct the final 3D model. Experiments on the DreamBooth dataset demonstrate that the proposed method achieves a CLIP-I score of 0.697 and a CLIP-T score of 0.295, significantly outperforming existing meth-ods including DreamBooth3D, MVDream, and GaussianEditor in terms of both geometry accuracy and visual fidelity, validating the effectiveness of the proposed approach.

     

/

返回文章
返回