高级检索

基于分布维持帧间噪声传播机制的文本引导的视频编辑

Distribution-Preserving Inter-frame Noise Propagation Mechanism for Text-guided Video Editing

  • 摘要: 近年来,大规模文生图扩散模型在文本引导的视频编辑领域展现出显著潜力。然而,由于缺乏对视频时序特性的精确建模,生成视频常面临严重的时序不一致问题。针对该挑战,提出一种基于分布保持帧间噪声传播机制的文本引导视频编辑方法。首先,利用DDIM反演算法将源视频转换为潜在空间的先验噪声;其次,在视频扩散模型去噪阶段的同一时间步内,对当前帧及其邻帧的预测噪声样本进行加权聚合,显式实现帧间信息的有效传播;再次,在聚合过程中引入严格的方差保持约束,确保传播后噪声样本在统计上服从原始高斯分布,避免破坏基线模型的去噪稳定性;最后,将该传播机制以无需额外训练的方式无缝集成至现有视频编辑基线模型,在目标文本提示引导下生成高质量编辑视频。在LOVEU-TGVE等数据集上的实验结果表明,所提方法在不引入额外计算开销的前提下,将基线模型的帧间时序一致性、文本忠实度和人类偏好评分分别提升2.61、1.74和0.28,并在多目标编辑及大规模文生视频基线模型上均展现出卓越性能。

     

    Abstract: Recently, large-scale text-to-image diffusion models have shown significant potential in text-guided video editing. However, due to the lack of precise temporal modeling, the generated videos often suffer from se-vere temporal inconsistencies. To address this challenge, we propose a novel text-guided video editing method based on a distribution-preserving inter-frame noise propagation mechanism. Specifically, we first employ DDIM inversion to map the source video into latent prior noise. Then, within the same timestep of the denoising phase, we perform a weighted aggregation of the predicted noise samples from the current and adjacent frames, explicitly enabling effective inter-frame information propagation. Furthermore, a strict variance-preserving constraint is introduced during the aggregation to ensure the propagated noise statistically conforms to the original Gaussian distribution, thereby avoiding the disruption of the baseline model's denoising stability. Finally, this training-free mechanism is seamlessly integrated into existing baselines to generate high-quality edited videos guided by target prompts. Extensive experiments on da-tasets such as LOVEU-TGVE demonstrate that, without introducing additional computational overhead, our method improves the temporal consistency, textual alignment, and human preference scores of baselines by 2.61, 1.74, and 0.28, respectively. Moreover, it exhibits superior performance in multi-object editing and generalizes well to large-scale text-to-video models.

     

/

返回文章
返回