Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

📄 arXiv: 2607.09581 📥 PDF

作者: Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Ruoshi Zhang, Yi Lu, Gang Cheng, Bang Zhang

分类: cs.CV, cs.SD

发布日期: 2026-07-20


💡 一句话要点

提出层次化框架以解决长时间音乐到舞蹈生成问题

🎯 匹配领域: 支柱三:空间感知与语义 (Perception & Semantics)

关键词: 音乐到舞蹈生成 层次化框架 时间一致性 动态帧率适应 光流损失函数 多模态生成 长视频合成

📋 核心要点

  1. 现有方法在生成超过20秒的舞蹈视频时,常常出现时间漂移和身份不一致等问题。
  2. 提出的层次化框架通过全局关键帧规划和局部时间细化,利用完整音乐上下文确保长时间一致性。
  3. 实验结果表明,该框架生成的720p/30fps视频超过一分钟,且在五种舞蹈风格中表现出优越的稳定性。

📝 摘要(中文)

生成长时间、高分辨率且节奏同步的舞蹈视频直接从音乐中提取仍然是一个重大挑战,尤其是当前扩散模型在20秒后表现不佳。现有方法无论是依赖中间3D骨架还是端到端视频合成,在延长时间时都面临时间漂移、身份不一致和重复运动模式等问题。为了解决这些限制,我们提出了一种新的层次化框架,用于分钟级一致的音乐到舞蹈生成。该方法将过程解耦为全局关键帧规划和局部时间细化,利用完整的音乐上下文确保长时间一致性。我们的框架在生成超过一分钟的720p/30fps视频时,展现出卓越的时间稳定性,且在五种不同舞蹈风格中表现出强大的通用性。

🔬 方法详解

问题定义:本论文旨在解决从音乐生成长时间舞蹈视频的挑战,现有方法在时间延续性和一致性方面存在显著不足,导致生成的视频质量下降。

核心思路:提出的层次化框架将生成过程分为全局关键帧规划和局部时间细化,充分利用音乐的全轨道上下文信息,以确保生成舞蹈的长时间一致性。

技术框架:整体架构包括两个主要模块:全局关键帧规划模块负责生成舞蹈的关键帧,而局部时间细化模块则对关键帧之间的过渡进行精细调整,确保动作的连贯性。

关键创新:本研究的创新点在于动态帧率适应技术,通过时间映射的RoPE嵌入实现精确对齐,以及基于光流的损失函数来增强运动的连续性,这些设计显著提升了生成视频的质量。

关键设计:在参数设置上,采用了动态帧率调整机制,并设计了光流损失函数以保持运动的自然流畅性,同时引入运动速度控制以确保快速动作中的细节保留。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,提出的框架成功生成720p/30fps的视频,时长超过一分钟,且在五种舞蹈风格中表现出优越的时间稳定性。与传统方法相比,生成视频的时间一致性和动作连贯性显著提升,标志着在一致性和长时段生成方面的突破。

🎯 应用场景

该研究的潜在应用领域包括舞蹈教育、娱乐产业以及虚拟现实中的舞蹈表演生成。通过自动化生成高质量舞蹈视频,可以为创作者提供新的创作工具,提升内容创作的效率和多样性。未来,该技术可能在个性化娱乐和互动体验中发挥重要作用。

📄 摘要(原文)

Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p/30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.