SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

📄 arXiv: 2607.21553v1 📥 PDF

作者: Junsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue, Haozhe Liu, Jingyu Xin, Yuyang Zhao, Tian Ye, Zhangjie Wu, Zian Wang, Daquan Zhou, Ping Luo, Song Han, Enze Xie

分类: cs.CV

发布日期: 2026-07-23

备注: 13 pages, 9 figures, 5 tables


💡 一句话要点

提出SANA-Video 2.0以高效生成720p视频

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)

关键词: 视频生成 混合注意力 深度学习 高效推理 长序列处理

📋 核心要点

  1. 现有视频生成方法在长序列处理和生成质量上存在性能瓶颈,尤其是在资源受限的环境中。
  2. SANA-Video 2.0通过混合线性注意力和注意力残差,提升了视频生成的效率和质量,能够在单个GPU上生成720p视频。
  3. 实验表明,SANA-Video 2.0在480p下以40步采样获得了84.30的VBench评分,且在720p/60s下比全软最大化基线快3.2倍。

📝 摘要(中文)

我们介绍了SANA-Video 2.0,这是一种在统一架构下实现的混合视频扩散变换器,规模为5B和14B。该模型旨在在单个GPU上生成高达720p的高质量视频,质量与全软最大化视频DiTs相当,同时保留线性注意力的长序列扩展性。通过结合门控线性注意力和周期性门控软最大化锚点,Hybrid Linear-Softmax Attention避免了二次注意力的使用,恢复了纯线性注意力所缺乏的全秩令牌交互。通过从头训练,SANA-Video 2.0直接学习完整的混合模型,优化了质量和效率的权衡。

🔬 方法详解

问题定义:论文旨在解决现有视频生成方法在长序列处理中的效率和质量问题,尤其是在计算资源有限的情况下,传统方法的二次注意力导致了性能瓶颈。

核心思路:SANA-Video 2.0采用混合线性注意力和注意力残差的设计,结合门控线性注意力和周期性门控软最大化锚点,旨在恢复全秩令牌交互,同时避免二次注意力的计算复杂度。

技术框架:该模型在统一架构下实现,包含多个模块,如门控线性注意力、周期性门控软最大化锚点和块注意力残差,能够有效传播信息并提升深层特征的有效秩。

关键创新:最重要的创新在于Hybrid Linear-Softmax Attention的引入,使得模型在保持高质量生成的同时,显著提高了计算效率,尤其是在长视频生成任务中。

关键设计:模型通过从头训练而非线性化预训练模型,采用了25%软最大化作为最佳质量效率权衡,并通过全栈Sol-Engine优化进一步加速了模型的推理速度。具体参数设置和损失函数设计在论文中有详细描述。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

SANA-Video 2.0在480p下以40步采样获得了84.30的VBench评分,且在720p/60s下的推理速度比全软最大化基线快3.2倍。经过全栈优化后,5B模型在720p/5s下的推理时间缩短至13.06秒,显示出120倍于Wan 2.2-A14B的速度提升。

🎯 应用场景

SANA-Video 2.0在视频生成领域具有广泛的应用潜力,尤其适用于需要高质量视频生成的场景,如影视制作、游戏开发和虚拟现实等。其高效的生成能力使得在资源受限的环境中也能实现高分辨率视频的实时生成,具有重要的实际价值和未来影响。

📄 摘要(原文)

We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.