高级检索

虚拟说话数字人视频生成方法综述

A Survey on Virtual Talking Digital Human Video Generation

  • 摘要: 面向高保真、强交互数字人的构建需求,系统梳理了说话数字人视频生成领域的关键技术演进。围绕形象生成、驱动机制、数据集构建、评估体系及代表性方法比较5个核心维度,展开系统性综述。在形象生成方面,梳理了二维图像生成、三维重建与实时渲染3条技术路线,阐明其在真实感、可控性及渲染效率等方面的特性差异;在驱动技术方面,归纳了以生成对抗网络与扩散模型为代表的二维驱动、基于三维参数化模型的驱动,以及基于神经辐射场与三维高斯溅射的驱动方法,并从输入模态、三维模型依赖、评估指标,以及驱动区域等7个维度进行对比分析,揭示了从驱动中间参数到直接生成像素再到三维场驱动的演进规律;在数据集层面,总结单人与多人数据集的构建特点,揭示了个性化保真与通用泛化之间的内在权衡;在评估体系方面,整合图像质量、视听同步及身份表情一致性3类9项指标,构建多维评估框架,指出现有指标在人类感知与实时性能方面仍存在不足;基于统一硬件平台与数据基准,对对抗网络、扩散模型、神经辐射场、三维高斯溅射4类路线的13项代表性方法进行评测,分析不同技术路线在实时性、生成质量与视听同步之间的权衡关系,并提出面向不同应用场景的技术选型建议。最后,展望了说话数字人在智能交互、高保真生成与实时部署方面的发展方向,汇总数据、效率、评估、场景适应性与伦理等方面的挑战及应对思路,以期为相关科研人员提供理论支撑与技术路线参考。

     

    Abstract: Aiming at the construction of high-fidelity and highly interactive digital humans, this paper systematically reviews the key technological evolution in talking digital human video generation around five core dimensions: avatar generation, driving mechanisms, dataset construction, evaluation frameworks, and comparative analysis of representative methods. For avatar generation, three technical routes—2D image generation, 3D reconstruction, and real-time rendering—are reviewed, elucidating their differences in photorealism, controllability, and rendering efficiency. For driving mechanisms, methods based on Generative Adversarial Networks (GANs), Diffusion Models (DMs), 3D parametric models, Neural Radiance Fields (NeRF), and 3D Gaussian Splatting (3DGS) are summarized and compared across seven dimensions including input modality, 3D model dependency, and controllable regions, revealing the evolutionary pattern from driving intermediate parameters to direct pixel generation and further to 3D field-based driving. At the dataset level, single-person and multi-person datasets are analyzed, exposing the trade-off between personalized fidelity and universal generalization. For evaluation, nine metrics across three categories—image quality, audio-visual synchronization, and identity-expression consistency—are integrated into a multi-dimensional framework, identifying deficiencies in human perceptual alignment and real-time assessment. Based on a unified hardware platform and data benchmark, thirteen representative methods across four technical routes are evaluated, analyzing trade-offs among real-time capability, generation quality, and audio-visual synchronization, with technology selection recommendations provided. Finally, future directions in intelligent interaction, high-fidelity generation, and real-time deployment are discussed, along with challenges in data, efficiency, evaluation, scene adaptability, and ethics, aiming to provide theoretical support and technical roadmap references for researchers.

     

/

返回文章
返回