Abstract:
Aiming at the construction of high-fidelity and highly interactive digital humans, this paper systematically reviews the key technological evolution in talking digital human video generation around five core dimensions: avatar generation, driving mechanisms, dataset construction, evaluation frameworks, and comparative analysis of representative methods. For avatar generation, three technical routes—2D image generation, 3D reconstruction, and real-time rendering—are reviewed, elucidating their differences in photorealism, controllability, and rendering efficiency. For driving mechanisms, methods based on Generative Adversarial Networks (GANs), Diffusion Models (DMs), 3D parametric models, Neural Radiance Fields (NeRF), and 3D Gaussian Splatting (3DGS) are summarized and compared across seven dimensions including input modality, 3D model dependency, and controllable regions, revealing the evolutionary pattern from driving intermediate parameters to direct pixel generation and further to 3D field-based driving. At the dataset level, single-person and multi-person datasets are analyzed, exposing the trade-off between personalized fidelity and universal generalization. For evaluation, nine metrics across three categories—image quality, audio-visual synchronization, and identity-expression consistency—are integrated into a multi-dimensional framework, identifying deficiencies in human perceptual alignment and real-time assessment. Based on a unified hardware platform and data benchmark, thirteen representative methods across four technical routes are evaluated, analyzing trade-offs among real-time capability, generation quality, and audio-visual synchronization, with technology selection recommendations provided. Finally, future directions in intelligent interaction, high-fidelity generation, and real-time deployment are discussed, along with challenges in data, efficiency, evaluation, scene adaptability, and ethics, aiming to provide theoretical support and technical roadmap references for researchers.